Gathering data about your cluster¶
You can gather debugging information about your OpenShift Container Platform cluster to provide to Red Hat Support when opening a support case.
About the must-gather tool¶
The oc adm must-gather CLI command collects the information from your cluster that is most likely needed for debugging issues, including:
- Resource definitions
- Service logs
By default, the oc adm must-gather command uses the default plugin image and writes into ./must-gather.local.
Alternatively, you can collect specific information by running the command with the appropriate arguments as described in the following sections:
-
To collect data related to one or more specific features, use the
--imageargument with an image, as listed in a following section.For example:
-
To collect the audit logs, use the
-- /usr/bin/gather_audit_logsargument, as described in a following section.
For example:
Note
- Audit logs are not collected as part of the default set of information to reduce the size of the files.
- On a Windows operating system, install the
cwRsyncclient and add to thePATHvariable for use with theoc rsynccommand.
When you run oc adm must-gather, a new pod with a random name is created in a new project on the cluster. The data is collected on that pod and saved in a new directory that starts with must-gather.local in the current working directory.
For example:
NAMESPACE NAME READY STATUS RESTARTS AGE
...
openshift-must-gather-5drcj must-gather-bklx4 2/2 Running 0 72s
openshift-must-gather-5drcj must-gather-s8sdh 2/2 Running 0 72s
...
Optionally, you can run the oc adm must-gather command in a specific namespace by using the --run-namespace option.
For example:
$ oc adm must-gather --run-namespace <namespace> \
--image=registry.redhat.io/container-native-virtualization/cnv-must-gather-rhel9:v4.22.6
Gathering data about your cluster for Red Hat Support¶
You can gather debugging information about your cluster by using the oc adm must-gather CLI command.
If you are gathering information to debug a self-managed hosted cluster, see "Gathering information to troubleshoot hosted control planes".
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - The OpenShift Container Platform CLI (
oc) is installed.
Procedure
-
Navigate to the directory where you want to store the
must-gatherdata.Note
If your cluster is in a disconnected environment, you must take additional steps. If your mirror registry has a trusted CA, you must first add the trusted CA to the cluster. For all clusters in disconnected environments, you must import the default
must-gatherimage as an image stream. -
Run the
oc adm must-gathercommand:Warning
If you are in a disconnected environment, use the
--imageflag as part of must-gather and point to the payload image.Note
Because this command picks a random control plane node by default, the pod might be scheduled to a control plane node that is in the
NotReadyandSchedulingDisabledstate.-
If this command fails, for example, if you cannot schedule a pod on your cluster, then use the
oc adm inspectcommand to gather information for particular resources.Note
Contact Red Hat Support for the recommended resources to gather.
-
-
Create a compressed file from the
must-gatherdirectory that was just created in your working directory. Make sure you provide the date and cluster ID for the unique must-gather data. For more information about how to find the cluster ID, see How to find the cluster-id or name on OpenShift cluster. For example, on a computer that uses a Linux operating system, run the following command:where:
<must_gather_local_dir>- Replace with the actual directory name.
-
Attach the compressed file to your support case on the the Customer Support page of the Red Hat Customer Portal.
Reducing the size of must-gather output¶
The oc adm must-gather command collects comprehensive cluster information. However, a full data collection can result in a large file that is difficult to upload and analyze and could result in timeouts.
To manage the output size and target your data collection for more effective troubleshooting, you can pass specific flags to the underlying gather script or scope the collection to particular resources.
Gathering data for specific resources¶
Instead of collecting data for the entire cluster, you can direct the must-gather tool to inspect a specific resource. This method is highly effective for isolating issues within a single project, Operator, or application.
The must-gather tool uses oc adm inspect internally. You can specify what to inspect by passing the inspect command and its arguments after the -- separator.
Procedure
-
To gather data for a specific namespace, such as
my-project, run the following command: -
This command collects all standard resources within the
my-projectnamespace, including logs from pods in that namespace, but excludes cluster-scoped resources. -
To gather data related to a specific Cluster Operator, such as
openshift-apiserver, run the following command: -
To exclude rotated logs, such as
*.gzor*.1files, from data collection, set theREDUCE_LOGSenvironment variable by running the following command: -
To exclude logs entirely and significantly reduce the size of the
must-gatherarchive, add a double dash (--) afteroc adm must-gathercommand and add the--no-logsargument:
Must-gather flags¶
The flags listed in the following table are available to use with the oc adm must-gather command.
OpenShift Container Platform flags for oc adm must-gather
| Flag | Example command | Description |
|---|---|---|
--all-images |
oc adm must-gather --all-images=false |
Collect must-gather data using the default image for all Operators on the cluster that are annotated with operators.openshift.io/must-gather-image. |
--dest-dir |
oc adm must-gather --dest-dir='<directory_name>' |
Set a specific directory on the local machine where the gathered data is written. |
--host-network |
oc adm must-gather --host-network=false |
Run must-gather pods as hostNetwork: true. Relevant if a specific command and image needs to capture host-level data. |
--image |
oc adm must-gather --image=[<plugin_image>] |
Specify a must-gather plugin image to run. If not specified, OpenShift Container Platform's default must-gather image is used. |
--image-stream |
oc adm must-gather --image-stream=[<image_stream>] |
Specify an<image_stream> using a namespace or name:tag value containing a must-gather plugin image to run. |
--node-name |
oc adm must-gather --node-name='<node>' |
Set a specific node to use. If not specified, by default a random master is used. |
--node-selector |
oc adm must-gather --node-selector='<node_selector_name>' |
Set a specific node selector to use. Only relevant when specifying a command and image which needs to capture data on a set of cluster nodes simultaneously. |
--run-namespace |
oc adm must-gather --run-namespace='<namespace>' |
An existing privileged namespace where must-gather pods should run. If not specified, a temporary namespace is generated. |
--since |
oc adm must-gather --since=<time> |
Only return logs newer than the specified duration. Defaults to all logs. Plugins are encouraged but not required to support this. Only one since-time or since may be used. |
--since-time |
oc adm must-gather --since-time='<date_and_time>' |
Only return logs after a specific date and time, expressed in (RFC3339) format. Defaults to all logs. Plugins are encouraged but not required to support this. Only one since-time or since may be used. |
--source-dir |
oc adm must-gather --source-dir='/<directory_name>/' |
Set the specific directory on the pod where you copy the gathered data from. |
--timeout |
oc adm must-gather --timeout='<time>' |
The length of time to gather data before timing out, expressed as seconds, minutes, or hours, for example, 3s, 5m, or 2h. Time specified must be higher than zero. Defaults to 10 minutes if not specified. |
--volume-percentage |
oc adm must-gather --volume-percentage=<percent> |
Specify maximum percentage of pod’s allocated volume that can be used for must-gather. If this limit is exceeded, must-gather stops gathering, but still copies gathered data. Defaults to 30% if not specified. |
Gathering data about specific features¶
You can gather debugging information about specific features by using the oc adm must-gather CLI command with the --image or --image-stream argument. The must-gather tool supports multiple images, so you can gather data about more than one feature by running a single command.
Supported must-gather images
| Image | Purpose |
|---|---|
registry.redhat.io/container-native-virtualization/cnv-must-gather-rhel9:v4.22.6 |
Data collection for OpenShift Virtualization. |
registry.redhat.io/openshift-serverless-1/svls-must-gather-rhel8 |
Data collection for OpenShift Serverless. |
registry.redhat.io/openshift-service-mesh/istio-must-gather-rhel8:<installed_version_service_mesh> |
Data collection for Red Hat OpenShift Service Mesh. |
registry.redhat.io/multicluster-engine/must-gather-rhel8 |
Data collection for hosted control planes. |
registry.redhat.io/odf4/odf-must-gather-rhel9:v<installed_version_ODF> |
Data collection for Red Hat OpenShift Data Foundation. |
registry.redhat.io/openshift-logging/cluster-logging-rhel9-operator:v<installed_version_logging> |
Data collection for logging. |
quay.io/netobserv/must-gather |
Data collection for the Network Observability Operator. |
registry.redhat.io/openshift4/ose-local-storage-mustgather-rhel9:v<installed_version_LSO> |
Data collection for Local Storage Operator. |
registry.redhat.io/openshift-sandboxed-containers/osc-must-gather-rhel8:v<installed_version_sandboxed_containers> |
Data collection for OpenShift sandboxed containers. |
registry.redhat.io/workload-availability/node-healthcheck-must-gather-rhel8:v<installed_version_NHC> |
Data collection for the Red Hat Workload Availability Operators, including the Self Node Remediation (SNR) Operator, the Fence Agents Remediation (FAR) Operator, the Machine Deletion Remediation (MDR) Operator, the Node Health Check (NHC) Operator, and the Node Maintenance Operator (NMO). Use this image if your NHC Operator version is earlier than 0.9.0. For more information, see the "Gathering data" section for the specific Operator in Remediation, fencing, and maintenance (Workload Availability for Red Hat OpenShift documentation). |
registry.redhat.io/workload-availability/node-healthcheck-must-gather-rhel9:v<installed_version_NHC> |
Data collection for the Red Hat Workload Availability Operators, including the Self Node Remediation (SNR) Operator, the Fence Agents Remediation (FAR) Operator, the Machine Deletion Remediation (MDR) Operator, the Node Health Check (NHC) Operator, and the Node Maintenance Operator (NMO). Use this image if your NHC Operator version is 0.9.0. or later. For more information, see the "Gathering data" section for the specific Operator in Remediation, fencing, and maintenance (Workload Availability for Red Hat OpenShift documentation). |
registry.redhat.io/openshift4/numaresources-must-gather-rhel9:v<installed-version-nro> |
Data collection for the NUMA Resources Operator (NRO). |
registry.redhat.io/openshift4/ptp-must-gather-rhel8:v<installed-version-ptp> |
Data collection for the PTP Operator. |
registry.redhat.io/openshift-gitops-1/must-gather-rhel8:v<installed_version_GitOps> |
Data collection for Red Hat OpenShift GitOps. |
registry.redhat.io/openshift4/ose-secrets-store-csi-mustgather-rhel9:v<installed_version_secret_store> |
Data collection for the Secrets Store CSI Driver Operator. |
registry.redhat.io/lvms4/lvms-must-gather-rhel9:v<installed_version_LVMS> |
Data collection for the LVM Operator. |
registry.redhat.io/compliance/openshift-compliance-must-gather-rhel8:<digest-version> |
Data collection for the Compliance Operator. |
Note
To determine the latest version for an OpenShift Container Platform component’s image, see the OpenShift Operator Life Cycles web page on the Red Hat Customer Portal.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - The OpenShift Container Platform CLI (
oc) is installed.
Procedure
-
Navigate to the directory where you want to store the
must-gatherdata. -
Run the
oc adm must-gathercommand with one or more--imageor--image-streamarguments.Note
- To collect the default
must-gatherdata in addition to specific feature data, add the--image-stream=openshift/must-gatherargument. - For information on gathering data about the Custom Metrics Autoscaler, see the Additional resources section that follows.
For example, the following command gathers both the default cluster data and information specific to OpenShift Virtualization:
$ oc adm must-gather \ --image-stream=openshift/must-gather \ --image=registry.redhat.io/container-native-virtualization/cnv-must-gather-rhel9:v4.22.6You can use the
must-gathertool with additional arguments to gather data that is specifically related to OpenShift Logging and the Red Hat OpenShift Logging Operator in your cluster. For OpenShift Logging, run the following command:$ oc adm must-gather --image=$(oc -n openshift-logging get deployment.apps/cluster-logging-operator \ -o jsonpath='{.spec.template.spec.containers[?(@.name == "cluster-logging-operator")].image}')Example must-gather output for OpenShift Logging├── cluster-logging │ ├── clo │ │ ├── cluster-logging-operator-74dd5994f-6ttgt │ │ ├── clusterlogforwarder_cr │ │ ├── cr │ │ ├── csv │ │ ├── deployment │ │ └── logforwarding_cr │ ├── collector │ │ ├── fluentd-2tr64 │ ├── eo │ │ ├── csv │ │ ├── deployment │ │ └── elasticsearch-operator-7dc7d97b9d-jb4r4 │ ├── es │ │ ├── cluster-elasticsearch │ │ │ ├── aliases │ │ │ ├── health │ │ │ ├── indices │ │ │ ├── latest_documents.json │ │ │ ├── nodes │ │ │ ├── nodes_stats.json │ │ │ └── thread_pool │ │ ├── cr │ │ ├── elasticsearch-cdm-lp8l38m0-1-794d6dd989-4jxms │ │ └── logs │ │ ├── elasticsearch-cdm-lp8l38m0-1-794d6dd989-4jxms │ ├── install │ │ ├── co_logs │ │ ├── install_plan │ │ ├── olmo_logs │ │ └── subscription │ └── kibana │ ├── cr │ ├── kibana-9d69668d4-2rkvz ├── cluster-scoped-resources │ └── core │ ├── nodes │ │ ├── ip-10-0-146-180.eu-west-1.compute.internal.yaml │ └── persistentvolumes │ ├── pvc-0a8d65d9-54aa-4c44-9ecc-33d9381e41c1.yaml ├── event-filter.html ├── gather-debug.log └── namespaces ├── openshift-logging │ ├── apps │ │ ├── daemonsets.yaml │ │ ├── deployments.yaml │ │ ├── replicasets.yaml │ │ └── statefulsets.yaml │ ├── batch │ │ ├── cronjobs.yaml │ │ └── jobs.yaml │ ├── core │ │ ├── configmaps.yaml │ │ ├── endpoints.yaml │ │ ├── events │ │ │ ├── elasticsearch-im-app-1596020400-gm6nl.1626341a296c16a1.yaml │ │ │ ├── elasticsearch-im-audit-1596020400-9l9n4.1626341a2af81bbd.yaml │ │ │ ├── elasticsearch-im-infra-1596020400-v98tk.1626341a2d821069.yaml │ │ │ ├── elasticsearch-im-app-1596020400-cc5vc.1626341a3019b238.yaml │ │ │ ├── elasticsearch-im-audit-1596020400-s8d5s.1626341a31f7b315.yaml │ │ │ ├── elasticsearch-im-infra-1596020400-7mgv8.1626341a35ea59ed.yaml │ │ ├── events.yaml │ │ ├── persistentvolumeclaims.yaml │ │ ├── pods.yaml │ │ ├── replicationcontrollers.yaml │ │ ├── secrets.yaml │ │ └── services.yaml │ ├── openshift-logging.yaml │ ├── pods │ │ ├── cluster-logging-operator-74dd5994f-6ttgt │ │ │ ├── cluster-logging-operator │ │ │ │ └── cluster-logging-operator │ │ │ │ └── logs │ │ │ │ ├── current.log │ │ │ │ ├── previous.insecure.log │ │ │ │ └── previous.log │ │ │ └── cluster-logging-operator-74dd5994f-6ttgt.yaml │ │ ├── cluster-logging-operator-registry-6df49d7d4-mxxff │ │ │ ├── cluster-logging-operator-registry │ │ │ │ └── cluster-logging-operator-registry │ │ │ │ └── logs │ │ │ │ ├── current.log │ │ │ │ ├── previous.insecure.log │ │ │ │ └── previous.log │ │ │ ├── cluster-logging-operator-registry-6df49d7d4-mxxff.yaml │ │ │ └── mutate-csv-and-generate-sqlite-db │ │ │ └── mutate-csv-and-generate-sqlite-db │ │ │ └── logs │ │ │ ├── current.log │ │ │ ├── previous.insecure.log │ │ │ └── previous.log │ │ ├── elasticsearch-cdm-lp8l38m0-1-794d6dd989-4jxms │ │ ├── elasticsearch-im-app-1596030300-bpgcx │ │ │ ├── elasticsearch-im-app-1596030300-bpgcx.yaml │ │ │ └── indexmanagement │ │ │ └── indexmanagement │ │ │ └── logs │ │ │ ├── current.log │ │ │ ├── previous.insecure.log │ │ │ └── previous.log │ │ ├── fluentd-2tr64 │ │ │ ├── fluentd │ │ │ │ └── fluentd │ │ │ │ └── logs │ │ │ │ ├── current.log │ │ │ │ ├── previous.insecure.log │ │ │ │ └── previous.log │ │ │ ├── fluentd-2tr64.yaml │ │ │ └── fluentd-init │ │ │ └── fluentd-init │ │ │ └── logs │ │ │ ├── current.log │ │ │ ├── previous.insecure.log │ │ │ └── previous.log │ │ ├── kibana-9d69668d4-2rkvz │ │ │ ├── kibana │ │ │ │ └── kibana │ │ │ │ └── logs │ │ │ │ ├── current.log │ │ │ │ ├── previous.insecure.log │ │ │ │ └── previous.log │ │ │ ├── kibana-9d69668d4-2rkvz.yaml │ │ │ └── kibana-proxy │ │ │ └── kibana-proxy │ │ │ └── logs │ │ │ ├── current.log │ │ │ ├── previous.insecure.log │ │ │ └── previous.log │ └── route.openshift.io │ └── routes.yaml └── openshift-operators-redhat ├── ... - To collect the default
-
Run the
oc adm must-gathercommand with one or more--imageor--image-streamarguments. For example, the following command gathers both the default cluster data and information specific to KubeVirt: -
Create a compressed file from the
must-gatherdirectory that was just created in your working directory. Make sure you provide the date and cluster ID for the unique must-gather data. For more information about how to find the cluster ID, see How to find the cluster-id or name on OpenShift cluster. For example, on a computer that uses a Linux operating system, run the following command:where:
<must_gather_local_dir>- Replace with the actual directory name.
-
Attach the compressed file to your support case on the the Customer Support page of the Red Hat Customer Portal.
Additional resources
- Gathering debugging data for the Custom Metrics Autoscaler
- Red Hat OpenShift Container Platform Life Cycle Policy
Gathering network logs¶
You can gather network logs on all nodes in a cluster.
Procedure
-
Run the
oc adm must-gathercommand with-- gather_network_logs:Note
By default, the
must-gathertool collects the OVNnbdbandsbdbdatabases from all of the nodes in the cluster. Adding the-- gather_network_logsoption to include additional logs that contain OVN-Kubernetes transactions for OVNnbdbdatabase. -
Create a compressed file from the
must-gatherdirectory that was just created in your working directory. Make sure you provide the date and cluster ID for the unique must-gather data. For more information about how to find the cluster ID, see How to find the cluster-id or name on OpenShift cluster. For example, on a computer that uses a Linux operating system, run the following command:Replace the
<must_gather_local_dir>placeholder with the actual directory name. -
Attach the compressed file to your support case on the the Customer Support page of the Red Hat Customer Portal.
Changing the must-gather storage limit¶
When using the oc adm must-gather command to collect data the default maximum storage for the information is 30% of the storage capacity of the container. After the 30% limit is reached the container is killed and the gathering process stops. Information already gathered is downloaded to your local storage. To run the must-gather command again, you need either a container with more storage capacity or to adjust the maximum volume percentage.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - The OpenShift CLI (
oc) is installed.
Procedure
-
Run the
oc adm must-gathercommand with thevolume-percentageflag. The new value cannot exceed 100.If the container reaches the storage limit, an error message similar to the following example is generated:
Self-service Technical Supportability Review¶
You can use the self-service Technical Supportability Review (TSR) on the Red Hat Customer Portal to validate your cluster configuration against Red Hat common practices.
Note
The must-gather tool collects diagnostic information about your cluster, including resource definitions, service logs, and configuration data. For more information, see "Gathering data about your cluster" in the OpenShift Container Platform documentation.
The self-service TSR uses AI to evaluate your cluster’s must-gather data and provides a prioritized executive summary of recommendations. This serves as a starting point to help you identify and resolve potential issues before they impact your environment.
The TSR performs hundreds of checks across the OpenShift Container Platform platform, including . Coverage is continually expanding.
When to use the self-service TSR tool¶
Integrating the self-service TSR into your regular operational workflow can be helpful in the following scenarios:
- Routine benchmarking
- Use the TSR quarterly to benchmark cluster health and plan for routine maintenance activities.
- Pre-flight checks
- Validate your cluster configuration before major structural changes, including upgrades, migrations, and expansions.
- Critical event preparation
- Confirm cluster stability ahead of high-traffic business events, such as seasonal peaks, or operational milestones, such as year-end shutdowns, business continuity drills, and compliance audits.
How to access the TSR¶
To run a self-service review, upload your cluster’s must-gather data to the Analyze tab in the Support section of the Red Hat Customer Portal. For a direct link, see "Technical Supportability Review with AI tool" in the Additional resources section. The Analyze feature generates a prioritized executive summary that identifies your cluster’s top risks and recommends corrective actions. Review the recommendations and implement the suggested corrective actions to address the identified risks.
The self-service TSR provides a solid baseline for cluster health. If you need additional guidance or a more comprehensive review, contact your Red Hat account team to arrange an assisted review through a Technical Account Manager (TAM) or Red Hat consultant. An assisted review includes human analysis, deeper coverage, and access to checks that are updated more frequently than the self-service version.
Additional resources
- Technical Supportability Review with AI tool
- Red Hat Technical Supportability Review with AI: Proactive AI-Driven Cluster Assessments for OpenShift Container Platform
About Support Log Gather¶
Support Log Gather Operator builds on the functionality of the traditional must-gather tool to automate the collection of debugging data. It streamlines troubleshooting by packaging the collected information into a single .tar file and automatically uploading it to the specified Red Hat Support case.
Warning
Support Log Gather is a Technology Preview feature only. Technology Preview features are not supported with Red Hat production service level agreements (SLAs) and might not be functionally complete. Red Hat does not recommend using them in production. These features provide early access to upcoming product features, enabling customers to test functionality and provide feedback during the development process.
For more information about the support scope of Red Hat Technology Preview features, see Technology Preview Features Support Scope.
The key features of Support Log Gather include the following:
- No administrator privileges required: Enables you to collect and upload logs without needing elevated permissions, making it easier for non-administrators to gather data securely.
- Simplified log collection: Collects debugging data from the cluster, such as resource definitions and service logs.
- Configurable data upload: Provides configuration options to either automatically upload the
.tarfile to a support case, or store it locally for manual upload.
Installing Support Log Gather by using the web console¶
You can use the web console to install the Support Log Gather.
Warning
Support Log Gather is a Technology Preview feature only. Technology Preview features are not supported with Red Hat production service level agreements (SLAs) and might not be functionally complete. Red Hat does not recommend using them in production. These features provide early access to upcoming product features, enabling customers to test functionality and provide feedback during the development process.
For more information about the support scope of Red Hat Technology Preview features, see Technology Preview Features Support Scope.
Prerequisites
- You have access to the cluster with
cluster-adminprivileges. - You have access to the OpenShift Container Platform web console.
Procedure
-
Log in to the OpenShift Container Platform web console.
-
Navigate to Ecosystem → Software Catalog.
-
In the filter box, enter Support Log Gather.
-
Select Support Log Gather.
-
From Version list, select the Support Log Gather version, and click Install.
-
On the Install Operator page, configure the installation settings:
-
Choose the Installed Namespace for the Operator.
The default Operator namespace is
must-gather-operator. Themust-gather-operatornamespace is created automatically if it does not exist. -
Select an Update approval strategy:
- Select Automatic to have the Operator Lifecycle Manager (OLM) update the Operator automatically when a newer version is available.
- Select Manual if Operator updates must be approved by a user with appropriate credentials.
-
Click Install.
-
Verification
-
Verify that the Operator is installed successfully:
- Navigate to Ecosystem → Software Catalog.
- Verify that Support Log Gather is listed with a Status of Succeeded in the
must-gather-operatornamespace.
-
Verify that Support Log Gather pods are running:
-
Navigate to Workloads → Pods
-
Verify that the status of the Support Log Gather pods is Running.
You can use the Support Log Gather only after the pods are up and running.
-
Installing Support Log Gather by using the CLI¶
To enable automated log collection for support cases, you can install Support Log Gather from the command-line interface (CLI).
Warning
Support Log Gather is a Technology Preview feature only. Technology Preview features are not supported with Red Hat production service level agreements (SLAs) and might not be functionally complete. Red Hat does not recommend using them in production. These features provide early access to upcoming product features, enabling customers to test functionality and provide feedback during the development process.
For more information about the support scope of Red Hat Technology Preview features, see Technology Preview Features Support Scope.
Prerequisites
- You have access to the cluster with
cluster-adminprivileges.
Procedure
-
Create a new project named
must-gather-operatorby running the following command: -
Create an
OperatorGroupobject:-
Create a YAML file, for example,
operatorGroup.yaml, that defines theOperatorGroupobject: -
Create the
OperatorGroupobject by running the following command:
-
-
Create a
Subscriptionobject:-
Create a YAML file, for example,
subscription.yaml, that defines theSubscriptionobject: -
Create the
Subscriptionobject by running the following command:
-
Verification
-
Verify the status of the pods in the Operator namespace by running the following command.
Example outputNAME READY STATUS RESTARTS AGE must-gather-operator-657fc74d64-2gg2w 1/1 Running 0 13mThe status of all the pods must be
Running. -
Verify that the subscription is created by running the following command:
-
Verify that the Operator is installed by running the following command:
Configuring a Support Log Gather instance¶
You must create a MustGather custom resource (CR) from the command-line interface (CLI) to automate the collection of diagnostic data from your cluster. This process also automatically uploads the data to a Red Hat Support case.
Warning
Support Log Gather is a Technology Preview feature only. Technology Preview features are not supported with Red Hat production service level agreements (SLAs) and might not be functionally complete. Red Hat does not recommend using them in production. These features provide early access to upcoming product features, enabling customers to test functionality and provide feedback during the development process.
For more information about the support scope of Red Hat Technology Preview features, see Technology Preview Features Support Scope.
Prerequisites
- You have installed the OpenShift CLI (
oc) tool. - You have installed Support Log Gather in your cluster.
- You have a Red Hat Support case ID.
- You have created a Kubernetes secret containing your Red Hat Customer Portal credentials. The secret must contain a username field and a password field.
- If you are using a custom image, you have configured an
ImageStreamresource in the Operator namespace that references an approved custom image URL. - You have created a service account. If you are using a custom image, you have created a service account with permissions to access the
ImageStreamresource.
Procedure
-
Create a YAML file for the
MustGatherCR, such assupport-log-gather.yaml, that contains the following configuration:Example support-log-gather.yamlapiVersion: operator.openshift.io/v1alpha1 kind: MustGather metadata: name: example-mg namespace: must-gather-operator spec: serviceAccountName: my-service-account gatherSpec: command: - "/usr/bin/custom-gather" args: - "--verbose" - "--subsystem=network" imageStreamRef: name: "network-debug-tools" tag: "v1.2" proxyConfig: httpProxy: "http://proxy.example.com:8080" httpsProxy: "https://proxy.example.com:8443" noProxy: ".example.com,localhost" mustGatherTimeout: "1h30m9s" uploadTarget: type: SFTP sftp: caseID: "04230315" caseManagementAccountSecretRef: name: mustgather-creds host: "sftp.access.redhat.com" retainResourcesOnCompletion: true storage: type: PersistentVolume persistentVolume: claim: name: mustgather-pvc subPath: must-gather-bundles/case-04230315For more information on the configuration parameters, see "Configuration parameters for MustGather custom resource".
-
Create the
MustGatherobject by running the following command:
Verification
-
Verify that the
MustGatherCR was created by running the following command: -
Verify the status of the pods in the Operator namespace by running the following command.
Example outputNAME READY STATUS RESTARTS AGE must-gather-operator-657fc74d64-2gg2w 1/1 Running 0 13m example-mg-gk8m8 2/2 Running 0 13sA new pod with a name based on the
MustGatherCR must be created. The status of all the pods must beRunning. -
To monitor the progress of the file upload, view the logs of the upload container in the job pod by running the following command:
When successful, the process must create an archive and upload it to the Red Hat Secure File Transfer Protocol (SFTP) server for the specified case.
Configurations for reducing the must-gather log size¶
Large must-gather logs can take a significant amount of time to upload to support cases and also consume considerable cluster storage. You can optimize the size of the collected diagnostic data by applying specific configurations to your MustGather custom resource (CR).
The following examples demonstrate different methods for reducing the must-gather log size:
Skipping rotated logs You can exclude older, rotated log files, such as *.gz or *.1 files, from the collection by setting the shell variable REDUCE_LOGS=skip_rotated_logs before running the gather script.
apiVersion: operator.openshift.io/v1alpha1
kind: MustGather
metadata:
name: full-mustgather
spec:
serviceAccountName: must-gather-operator
gatherSpec:
command:
- /bin/sh
- -c
- |
REDUCE_LOGS=skip_rotated_logs gather
uploadTarget:
type: SFTP
sftp:
caseID: '02527285'
caseManagementAccountSecretRef:
name: sftp-access-rh-creds
internalUser: true
REDUCE_LOGS=skip_rotated_logs gather- Sets the
REDUCE_LOGSshell variable and executes thegatherscript. As a result, the script excludes the collection of rotated log files.
Additional resources
Configuration parameters for MustGather custom resource¶
You can manage your MustGather custom resource (CR) by creating a YAML file that specifies the parameters for data collection and the upload process. The following table provides an overview of the parameters that you can configure in the MustGather CR.
| Parameter name | Description | Type |
|---|---|---|
spec.gatherSpec.args |
Optional: Specifies a list of command-line arguments. The Operator passes this value to the args field of the container. If you do not specify spec.gatherSpec.command, the specified arguments are appended to the default command of the Operator. |
List of strings |
spec.gatherSpec.audit |
Optional: Specifies whether to collect audit logs. The valid values are true and false. You must not set this field if you are using a custom image, or spec.gatherSpec.command with the default image. |
boolean |
spec.gatherSpec.command |
Optional: Overrides the default command of the container. The Operator passes this value to the command field of the container. |
List of strings |
spec.gatherSpec.since |
Optional: Specifies a time duration to restrict log collection to entries newer than the specified duration. By default, the controller collects all available logs. You can specify either spec.gatherSpec.since or spec.gatherSpec.sinceTime, but not both. |
The value must be a number with a time unit. The valid units are s (seconds), m (minutes), or h (hours). |
spec.gatherSpec.sinceTime |
Optional: Specifies a timestamp to restrict log collection to entries newer than the specified timestamp. By default, the controller collects all available logs. You can specify either spec.gatherSpec.since or spec.gatherSpec.sinceTime, but not both. |
The value must be in RFC3339 format. |
spec.imageStreamRef |
Optional: Overrides the default image by defining a specific custom image. Note Each |
object |
spec.imageStreamRef.name |
Specifies the name of the ImageStream resource in the Operator namespace. |
string |
spec.imageStreamRef.tag |
Specifies the name of the tag within the ImageStream resource. |
string |
spec.mustGatherTimeout |
Optional: Specifies the time limit for the must-gather command to complete. |
The value must be a number with a time unit. The valid units are s (seconds), m (minutes), or h (hours). By default, no time limit is set. |
spec.retainResourcesOnCompletion |
Optional: Specifies whether to retain the must-gather job and its related resources after the completion of data collection. The valid values are true and false. The default value is false. |
boolean |
spec.serviceAccountName |
Optional: Specifies the name of the service account. The default value is default.Note Because the |
string |
spec.storage |
Optional: Defines the storage configuration for the must-gather bundle. |
Object |
spec.storage.persistentVolume |
Defines the details of the persistent volume. | Object |
spec.storage.persistentVolume.claim |
Defines the details of the persistent volume claim (PVC). | Object |
spec.storage.persistentVolume.claim.name |
Specifies the name of the PVC to be used for storage. | string |
spec.storage.persistentVolume.subPath |
Optional: Specifies the path within the PVC to store the bundle. | string |
spec.storage.type |
Defines the type of storage. The only supported value is PersistentVolume. |
string |
spec.uploadTarget |
Optional: Defines the upload location for the must-gather bundle. |
Object |
spec.uploadTarget.sftp.caseID |
Specifies the Red Hat Support case ID for which the diagnostic data is collected. | string |
spec.uploadTarget.sftp.caseManagementAccountSecretRef |
Defines the credentials required for authenticating and uploading the files to the Red Hat Customer Portal support case. The value must contain a username and password field. |
Object |
spec.uploadTarget.sftp.caseManagementAccountSecretRef.name |
Specifies the name of the Kubernetes secret that contains the credentials. | string |
spec.uploadTarget.sftp.host |
Optional: Specifies the destination server for the bundle upload. By default, the bundle is uploaded to sftp.access.redhat.com. |
|
spec.uploadTarget.sftp.internalUser |
Optional: Specifies whether the user provided in the caseManagementAccountSecretRef is a Red Hat internal user. The valid values are true and false. The default value is false. |
boolean |
spec.uploadTarget.type |
Specifies the type of upload location for the must-gather bundle. The only supported value is SFTP. |
string |
Note
If you do not specify spec.uploadTarget or spec.storage, the pod saves the data to an ephemeral volume and the data is permanently deleted when the pod terminates.
Uninstalling Support Log Gather¶
You can uninstall the Support Log Gather by using the web console.
Prerequisites
- You have access to the cluster with
cluster-adminprivileges. - You have access to the OpenShift Container Platform web console.
- The Support Log Gather is installed.
Procedure
-
Log in to the OpenShift Container Platform web console.
-
Uninstall the Support Log Gather Operator.
- Navigate to Ecosystem → Installed Operators.
- Click the Options menu
next to the Support Log Gather entry and click Uninstall Operator. - In the confirmation dialog, click Uninstall.
Removing Support Log Gather resources¶
Once you have uninstalled the Support Log Gather, you can remove the associated resources from your cluster.
Prerequisites
- You have access to the cluster with
cluster-adminprivileges. - You have access to the OpenShift Container Platform web console.
Procedure
-
Log in to the OpenShift Container Platform web console.
-
Delete the component deployments in the must-gather-operator namespace.:
-
Click the Project drop-down menu to view the list of all available projects, and select the must-gather-operator project.
-
Navigate to Workloads → Deployments.
-
Select the deployment that you want to delete.
-
Click the Actions drop-down menu, and select Delete Deployment.
-
In the confirmation dialog box, click Delete to delete the deployment.
-
Alternatively, delete deployments of the components present in the
must-gather-operatornamespace by using the command-line interface (CLI).
-
-
Optional: Remove the custom resource definitions (CRDs) that were installed by the Support Log Gather:
-
Navigate to Administration → CustomResourceDefinitions.
-
Enter
MustGatherin the Name field to filter the CRDs. -
Click the Options menu
next to each of the following CRDs, and select Delete Custom Resource Definition:MustGather
-
-
Optional: Remove the
must-gather-operatornamespace.- Navigate to Administration → Namespaces.
- Click the Options menu
next to the must-gather-operator and select Delete Namespace. - In the confirmation dialog box, enter
must-gather-operatorand click Delete.
Obtaining your cluster ID¶
When providing information to Red Hat Support, it is helpful to provide the unique identifier for your cluster. You can have your cluster ID autofilled by using the OpenShift Container Platform web console. You can also manually obtain your cluster ID by using the web console or the OpenShift CLI (oc).
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - You have access to the web console or the OpenShift CLI (
oc) installed.
Procedure
-
To open a support case and have your cluster ID autofilled using the web console:
- From the toolbar, navigate to (?) Help and select Share Feedback from the list.
- Click Open a support case from the Tell us about your experience window.
-
To manually obtain your cluster ID using the web console:
- Navigate to Home → Overview.
- The value is available in the Cluster ID field of the Details section.
-
To obtain your cluster ID using the OpenShift CLI (
oc), run the following command:
About sosreport¶
sosreport is a tool that collects configuration details, system information, and diagnostic data from Red Hat Enterprise Linux (RHEL) and Red Hat Enterprise Linux CoreOS (RHCOS) systems. sosreport provides a standardized way to collect diagnostic information relating to a node, which can then be provided to Red Hat Support for issue diagnosis.
In some support interactions, Red Hat Support may ask you to collect a sosreport archive for a specific OpenShift Container Platform node. For example, it might sometimes be necessary to review system logs or other node-specific data that is not included within the output of oc adm must-gather.
Generating a sosreport archive for an OpenShift Container Platform cluster node¶
The recommended way to generate a sosreport for an OpenShift Container Platform 4.22 cluster node is through a debug pod.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - You have SSH access to your hosts.
- You have installed the OpenShift CLI (
oc). - You have a Red Hat standard or premium Subscription.
- You have a Red Hat Customer Portal account.
- You have an existing Red Hat Support case ID.
Procedure
-
Obtain a list of cluster nodes:
-
Enter into a debug session on the target node. This step instantiates a debug pod called
<node_name>-debug:To enter into a debug session on the target node that is tainted with the
NoExecuteeffect, add a toleration to a dummy namespace, and start the debug pod in the dummy namespace: -
Set
/hostas the root directory within the debug shell. The debug pod mounts the host’s root file system in/hostwithin the pod. By changing the root directory to/host, you can run binaries contained in the host’s executable paths:Note
OpenShift Container Platform 4.22 cluster nodes running Red Hat Enterprise Linux CoreOS (RHCOS) are immutable and rely on Operators to apply cluster changes. Accessing cluster nodes by using SSH is not recommended. However, if the OpenShift Container Platform API is not available, or the kubelet is not properly functioning on the target node,
ocoperations will be impacted. In such situations, it is possible to access nodes usingssh core@<node>.<cluster_name>.<base_domain>instead. -
Start a
toolboxcontainer, which includes the required binaries and plugins to runsosreport:Note
If an existing
toolboxpod is already running, thetoolboxcommand outputs'toolbox-' already exists. Trying to start.... Remove the running toolbox container withpodman rm toolbox-and spawn a new toolbox container, to avoid issues withsosreportplugins. -
Collect a
sosreportarchive.-
Run the
sos reportcommand to collect necessary troubleshooting data oncrioandpodman:- where
-
-kenables you to definesosreportplugin parameters outside of the defaults.
-
Optional: To include information on OVN-Kubernetes networking configurations from a node in your report, run the following command:
-
Press Enter when prompted, to continue.
-
Provide the Red Hat Support case ID.
sosreportadds the ID to the archive’s file name. -
The
sosreportoutput provides the archive’s location and checksum. The following sample output references support case ID01234567:Your sosreport has been generated and saved in: /host/var/tmp/sosreport-my-cluster-node-01234567-2020-05-28-eyjknxt.tar.xz The checksum is: 382ffc167510fd71b4f12a4f40b97a4e- where
-
- The
sosreportarchive’s file path is outside of thechrootenvironment because the toolbox container mounts the host’s root directory at/host.
- The
-
-
Provide the
sosreportarchive to Red Hat Support for analysis, using one of the following methods.-
Upload the file to an existing Red Hat support case.
-
Concatenate the
sosreportarchive by running theoc debug node/<node_name>command and redirect the output to a file. This command assumes you have exited the previousoc debugsession:$ oc debug node/my-cluster-node -- bash -c 'cat /host/var/tmp/sosreport-my-cluster-node-01234567-2020-05-28-eyjknxt.tar.xz' > /tmp/sosreport-my-cluster-node-01234567-2020-05-28-eyjknxt.tar.xz- where
-
- The debug container mounts the host’s root directory at
/host. Reference the absolute path from the debug container’s root directory, including/host, when specifying target files for concatenation.
- The debug container mounts the host’s root directory at
Note
OpenShift Container Platform 4.22 cluster nodes running Red Hat Enterprise Linux CoreOS (RHCOS) are immutable and rely on Operators to apply cluster changes. Transferring a
sosreportarchive from a cluster node by usingscpis not recommended. However, if the OpenShift Container Platform API is not available, or the kubelet is not properly functioning on the target node,ocoperations will be impacted. In such situations, it is possible to copy asosreportarchive from a node by runningscp core@<node>.<cluster_name>.<base_domain>:<file_path> <local_path>. -
-
Querying bootstrap node journal logs¶
If you experience bootstrap-related issues, you can gather bootkube.service journald unit logs and container logs from the bootstrap node.
Prerequisites
- You have SSH access to your bootstrap node.
- You have the fully qualified domain name of the bootstrap node.
Procedure
-
Query
bootkube.servicejournaldunit logs from a bootstrap node during OpenShift Container Platform installation. Replace<bootstrap_fqdn>with the bootstrap node’s fully qualified domain name:Note
The
bootkube.servicelog on the bootstrap node outputs etcdconnection refusederrors, indicating that the bootstrap server is unable to connect to etcd on control plane nodes. After etcd has started on each control plane node and the nodes have joined the cluster, the errors should stop. -
Collect logs from the bootstrap node containers using
podmanon the bootstrap node. Replace<bootstrap_fqdn>with the bootstrap node’s fully qualified domain name:
Querying cluster node journal logs¶
You can gather journald unit logs and other logs within /var/log on individual cluster nodes.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - You have installed the OpenShift CLI (
oc). - Your API service is still functional.
- You have SSH access to your hosts.
Procedure
-
Query
kubeletjournaldunit logs from OpenShift Container Platform cluster nodes. The following example queries control plane nodes only:kubelet- Replace as appropriate to query other unit logs.
-
Collect logs from specific subdirectories under
/var/log/on cluster nodes.-
Retrieve a list of logs contained within a
/var/log/subdirectory. The following example lists files in/var/log/openshift-apiserver/on all control plane nodes: -
Inspect a specific log within a
/var/log/subdirectory. The following example outputs/var/log/openshift-apiserver/audit.logcontents from all control plane nodes: -
If the API is not functional, review the logs on each node using SSH instead. The following example tails
/var/log/openshift-apiserver/audit.log:$ ssh core@<master-node>.<cluster_name>.<base_domain> sudo tail -f /var/log/openshift-apiserver/audit.logNote
OpenShift Container Platform 4.22 cluster nodes running Red Hat Enterprise Linux CoreOS (RHCOS) are immutable and rely on Operators to apply cluster changes. Accessing cluster nodes by using SSH is not recommended. Before attempting to collect diagnostic data over SSH, review whether the data collected by running
oc adm must gatherand otheroccommands is sufficient instead. However, if the OpenShift Container Platform API is not available, or the kubelet is not properly functioning on the target node,ocoperations will be impacted. In such situations, it is possible to access nodes usingssh core@<node>.<cluster_name>.<base_domain>.
-
Network trace methods¶
Collecting network traces, in the form of packet capture records, can assist Red Hat Support with troubleshooting network issues.
OpenShift Container Platform supports two ways of performing a network trace. Review the following table and choose the method that meets your needs.
Supported methods of collecting a network trace
| Method | Benefits and capabilities |
|---|---|
| Collecting a host network trace | You perform a packet capture for a duration that you specify on one or more nodes at the same time. The packet capture files are transferred from nodes to the client machine when the specified duration is met. You can troubleshoot why a specific action triggers network communication issues. Run the packet capture, perform the action that triggers the issue, and use the logs to diagnose the issue. |
| Collecting a network trace from an OpenShift Container Platform node or container | You perform a packet capture on one node or one container. You run the tcpdump command interactively, so you can control the duration of the packet capture.You can start the packet capture manually, trigger the network communication issue, and then stop the packet capture manually. This method uses the cat command and shell redirection to copy the packet capture data from the node or container to the client machine. |
Collecting a host network trace¶
Sometimes, troubleshooting a network-related issue is simplified by tracing network communication and capturing packets on multiple nodes at the same time.
You can use a combination of the oc adm must-gather command and the registry.redhat.io/openshift4/network-tools-rhel8 container image to gather packet captures from nodes. Analyzing packet captures can help you troubleshoot network communication issues.
The oc adm must-gather command is used to run the tcpdump command in pods on specific nodes. The tcpdump command records the packet captures in the pods. When the tcpdump command exits, the oc adm must-gather command transfers the files with the packet captures from the pods to your client machine.
Tip
The sample command in the following procedure demonstrates performing a packet capture with the tcpdump command. However, you can run any command in the container image that is specified in the --image argument to gather troubleshooting information from multiple nodes at the same time.
Prerequisites
- You are logged in to OpenShift Container Platform as a user with the
cluster-adminrole. - You have installed the OpenShift CLI (
oc).
Procedure
-
Run a packet capture from the host network on some nodes by running the following command:
$ oc adm must-gather \ --dest-dir /tmp/captures \ --source-dir '/tmp/tcpdump/' \ --image registry.redhat.io/openshift4/network-tools-rhel8:latest \ --node-selector 'node-role.kubernetes.io/worker' \ --host-network=true \ --timeout 30s \ -- \ tcpdump -i any \ -w /tmp/tcpdump/%Y-%m-%dT%H:%M:%S.pcap -W 1 -G 300where:
--dest-dir /tmp/captures- The
--dest-dirargument specifies thatoc adm must-gatherstores the packet captures in directories that are relative to/tmp/captureson the client machine. You can specify any writable directory. --source-dir '/tmp/tcpdump/'- When
tcpdumpis run in the debug pod thatoc adm must-gatherstarts, the--source-dirargument specifies that the packet captures are temporarily stored in the/tmp/tcpdumpdirectory on the pod. --image registry.redhat.io/openshift4/network-tools-rhel8:latest- The
--imageargument specifies a container image that includes thetcpdumpcommand. --node-selector 'node-role.kubernetes.io/worker'- The
--node-selectorargument and example value specifies to perform the packet captures on the worker nodes. As an alternative, you can specify the--node-nameargument instead to run the packet capture on a single node. If you omit both the--node-selectorand the--node-nameargument, the packet captures are performed on all nodes. --host-network=true- The
--host-network=trueargument is required so that the packet captures are performed on the network interfaces of the node. --timeout 30s- The
--timeoutargument and value specify to run the debug pod for 30 seconds. If you do not specify the--timeoutargument and a duration, the debug pod runs for 10 minutes. -i any- The
-i anyargument for thetcpdumpcommand specifies to capture packets on all network interfaces. As an alternative, you can specify a network interface name.
-
Perform the action, such as accessing a web application, that triggers the network communication issue while the network trace captures packets.
-
Review the packet capture files that
oc adm must-gathertransferred from the pods to your client machine:tmp/captures ├── event-filter.html ├── ip-10-0-192-217-ec2-internal │ └── registry-redhat-io-openshift4-network-tools-rhel8-sha256-bca... │ └── 2022-01-13T19:31:31.pcap ├── ip-10-0-201-178-ec2-internal │ └── registry-redhat-io-openshift4-network-tools-rhel8-sha256-bca... │ └── 2022-01-13T19:31:30.pcap ├── ip-... └── timestampwhere:
ip-10-0-192-217-ec2-internal,ip-10-0-201-178-ec2-internal- The packet captures are stored in directories that identify the hostname, container, and file name. If you did not specify the
--node-selectorargument, then the directory level for the hostname is not present.
Collecting a network trace from an OpenShift Container Platform node or container¶
When investigating potential network-related OpenShift Container Platform issues, Red Hat Support might request a network packet trace from a specific OpenShift Container Platform cluster node or from a specific container. The recommended method to capture a network trace in OpenShift Container Platform is through a debug pod.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - You have installed the OpenShift CLI (
oc). - You have an existing Red Hat Support case ID.
- You have a Red Hat standard or premium Subscription.
- You have a Red Hat Customer Portal account.
- You have SSH access to your hosts.
Procedure
-
Obtain a list of cluster nodes:
-
Enter into a debug session on the target node. This step instantiates a debug pod called
<node_name>-debug: -
Set
/hostas the root directory within the debug shell. The debug pod mounts the host’s root file system in/hostwithin the pod. By changing the root directory to/host, you can run binaries contained in the host’s executable paths:Note
OpenShift Container Platform 4.22 cluster nodes running Red Hat Enterprise Linux CoreOS (RHCOS) are immutable and rely on Operators to apply cluster changes. Accessing cluster nodes by using SSH is not recommended. However, if the OpenShift Container Platform API is not available, or the kubelet is not properly functioning on the target node,
ocoperations will be impacted. In such situations, it is possible to access nodes usingssh core@<node>.<cluster_name>.<base_domain>instead. -
From within the
chrootenvironment console, obtain the node’s interface names: -
Start a
toolboxcontainer, which includes the required binaries and plugins to runsosreport:Note
If an existing
toolboxpod is already running, thetoolboxcommand outputs'toolbox-' already exists. Trying to start.... To avoidtcpdumpissues, remove the running toolbox container withpodman rm toolbox-and spawn a new toolbox container. -
Initiate a
tcpdumpsession on the cluster node and redirect output to a capture file. This example usesens5as the interface name:where:
/host/var/tmp/my-cluster-node_$(date +%d_%m_%Y-%H_%M_%S-%Z).pcap- The
tcpdumpcapture file’s path is outside of thechrootenvironment because the toolbox container mounts the host’s root directory at/host.
-
If a
tcpdumpcapture is required for a specific container on the node, follow these steps.-
Determine the target container ID. The
chroot hostcommand precedes thecrictlcommand in this step because the toolbox container mounts the host’s root directory at/host: -
Determine the container’s process ID. In this example, the container ID is
a7fe32346b120: -
Initiate a
tcpdumpsession on the container and redirect output to a capture file. This example uses49628as the container’s process ID andens5as the interface name. Thensentercommand enters the namespace of a target process and runs a command in its namespace. because the target process in this example is a container’s process ID, thetcpdumpcommand is run in the container’s namespace from the host:# nsenter -n -t 49628 -- tcpdump -nn -i ens5 -w /host/var/tmp/my-cluster-node-my-container_$(date +%d_%m_%Y-%H_%M_%S-%Z).pcapwhere:
/host/var/tmp/my-cluster-node-my-container_$(date +%d_%m_%Y-%H_%M_%S-%Z).pcap- The
tcpdumpcapture file’s path is outside of thechrootenvironment because the toolbox container mounts the host’s root directory at/host.
-
-
Provide the
tcpdumpcapture file to Red Hat Support for analysis, using one of the following methods.-
Upload the file to an existing Red Hat support case.
-
Concatenate the
sosreportarchive by running theoc debug node/<node_name>command and redirect the output to a file. This command assumes you have exited the previousoc debugsession:$ oc debug node/my-cluster-node -- bash -c 'cat /host/var/tmp/my-tcpdump-capture-file.pcap' > /tmp/my-tcpdump-capture-file.pcapwhere:
/host/var/tmp/my-tcpdump-capture-file.pcap- The debug container mounts the host’s root directory at
/host. Reference the absolute path from the debug container’s root directory, including/host, when specifying target files for concatenation.
Note
OpenShift Container Platform 4.22 cluster nodes running Red Hat Enterprise Linux CoreOS (RHCOS) are immutable and rely on Operators to apply cluster changes. Transferring a
tcpdumpcapture file from a cluster node by usingscpis not recommended. However, if the OpenShift Container Platform API is not available, or the kubelet is not properly functioning on the target node,ocoperations will be impacted. In such situations, it is possible to copy atcpdumpcapture file from a node by runningscp core@<node>.<cluster_name>.<base_domain>:<file_path> <local_path>. -
Navigate to an existing support case within the Customer Support page of the Red Hat Customer Portal.
-
Select Attach files and follow the prompts to upload the file.
-
-
Providing diagnostic data to Red Hat Support¶
When investigating OpenShift Container Platform issues, Red Hat Support might ask you to upload diagnostic data to a support case. Files can be uploaded to a support case through the Red Hat Customer Portal.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - You have installed the OpenShift CLI (
oc). - You have SSH access to your hosts.
- You have a Red Hat standard or premium Subscription.
- You have a Red Hat Customer Portal account.
- You have an existing Red Hat Support case ID.
Procedure
-
Upload diagnostic data to an existing Red Hat support case through the Red Hat Customer Portal.
-
Concatenate a diagnostic file contained on an OpenShift Container Platform node by using the
oc debug node/<node_name>command and redirect the output to a file. The following example copies/host/var/tmp/my-diagnostic-data.tar.gzfrom a debug container to/var/tmp/my-diagnostic-data.tar.gz:$ oc debug node/my-cluster-node -- bash -c 'cat /host/var/tmp/my-diagnostic-data.tar.gz' > /var/tmp/my-diagnostic-data.tar.gzwhere:
/host/var/tmp/my-diagnostic-data.tar.gz- The debug container mounts the host’s root directory at
/host. Reference the absolute path from the debug container’s root directory, including/host, when specifying target files for concatenation.
Note
OpenShift Container Platform 4.22 cluster nodes running Red Hat Enterprise Linux CoreOS (RHCOS) are immutable and rely on Operators to apply cluster changes. Transferring files from a cluster node by using
scpis not recommended. However, if the OpenShift Container Platform API is not available, or the kubelet is not properly functioning on the target node,ocoperations will be impacted. In such situations, it is possible to copy diagnostic files from a node by runningscp core@<node>.<cluster_name>.<base_domain>:<file_path> <local_path>. -
Navigate to an existing support case within the Customer Support page of the Red Hat Customer Portal.
-
Select Attach files and follow the prompts to upload the file.
-
About toolbox¶
toolbox is a tool that starts a container on a Red Hat Enterprise Linux CoreOS (RHCOS) system. The tool is primarily used to start a container that includes the required binaries and plugins that are needed to run commands such as sosreport.
The primary purpose for a toolbox container is to gather diagnostic information and to provide it to Red Hat Support. However, if additional diagnostic tools are required, you can add RPM packages or run an image that is an alternative to the standard support tools image.
Installing packages to a toolbox container¶
By default, running the toolbox command starts a container with the registry.redhat.io/rhel9/support-tools:latest image. This image contains the most frequently used support tools. If you need to collect node-specific data that requires a support tool that is not part of the image, you can install additional packages.
Prerequisites
- You have accessed a node with the
oc debug node/<node_name>command. - You can access your system as a user with root privileges.
Procedure
-
Set
/hostas the root directory within the debug shell. The debug pod mounts the host’s root file system in/hostwithin the pod. By changing the root directory to/host, you can run binaries contained in the host’s executable paths: -
Start the toolbox container:
-
Install the additional package, such as
wget:
Starting an alternative image with toolbox¶
By default, running the toolbox command starts a container with the registry.redhat.io/rhel9/support-tools:latest image.
Note
You can start an alternative image by creating a .toolboxrc file and specifying the image to run. However, running an older version of the support-tools image, such as registry.redhat.io/rhel8/support-tools:latest, is not supported on OpenShift Container Platform 4.22.
Prerequisites
- You have accessed a node with the
oc debug node/<node_name>command. - You can access your system as a user with root privileges.
Procedure
-
Set
/hostas the root directory within the debug shell. The debug pod mounts the host’s root file system in/hostwithin the pod. By changing the root directory to/host, you can run binaries contained in the host’s executable paths: -
Optional: If you need to use an alternative image instead of the default image, create a
.toolboxrcfile in the home directory for the root user ID, and specify the image metadata:where:
REGISTRY=quay.io- Optional: Specify an alternative container registry.
IMAGE=fedora/fedora:latest- Specify an alternative image to start.
TOOLBOX_NAME=toolbox-fedora-latest- Optional: Specify an alternative name for the toolbox container.
-
Start a toolbox container by entering the following command:
Note
If an existing
toolboxpod is already running, thetoolboxcommand outputs'toolbox-' already exists. Trying to start.... To avoid issues withsosreportplugins, remove the running toolbox container withpodman rm toolbox-and then spawn a new toolbox container.