Troubleshooting hosted control planes¶
If you encounter an issue with hosted control planes, you can gather information about the hosted cluster, OpenShift Container Platform, or other components so that you can determine the root cause and take steps to resolve it.
Gathering information to troubleshoot hosted control planes¶
When you need to troubleshoot an issue with hosted clusters, you can gather information by running the must-gather command. The command generates output for the management cluster and the hosted cluster.
The output for the management cluster contains the following content:
- Cluster-scoped resources: These resources are node definitions of the management cluster.
- The
hypershift-dumpcompressed file: This file is useful if you need to share the content with other people. - Namespaced resources: These resources include all of the objects from the relevant namespaces, such as config maps, services, events, and logs.
- Network logs: These logs include the OVN northbound and southbound databases and the status for each one.
- Hosted clusters: This level of output involves all of the resources inside of the hosted cluster.
The output for the hosted cluster contains the following content:
- Cluster-scoped resources: These resources include all of the cluster-wide objects, such as nodes and CRDs.
- Namespaced resources: These resources include all of the objects from the relevant namespaces, such as config maps, services, events, and logs.
Although the output does not contain any secret objects from the cluster, it can contain references to the names of secrets.
Prerequisites
- You must have
cluster-adminaccess to the management cluster. - You need the
namevalue for theHostedClusterresource and the namespace where the CR is deployed. - You must have the
hcpcommand-line interface installed. For more information, see "Installing the hosted control planes command-line interface". - You must have the OpenShift CLI (
oc) installed. - You must ensure that the
kubeconfigfile is loaded and is pointing to the management cluster.
Procedure
-
To gather the output for troubleshooting, enter the following command:
$ oc adm must-gather \ --image=registry.redhat.io/rhacm2/acm-must-gather-rhel9:v2.17 \ /usr/bin/gather hosted-cluster-namespace=HOSTEDCLUSTERNAMESPACE \ hosted-cluster-name=HOSTEDCLUSTERNAME \ --dest-dir=NAME ; tar -cvzf NAME.tgz NAMEwhere:
- The
hosted-cluster-namespace=HOSTEDCLUSTERNAMESPACEparameter is optional. If you do not include it, the command runs as though the hosted cluster is in the default namespace, which isclusters. - If you want to save the results of the command to a compressed file, specify the
--dest-dir=NAMEparameter and replaceNAMEwith the name of the directory where you want to save the results.
- The
Additional resources
Gathering OpenShift Container Platform data for a hosted cluster¶
You can gather OpenShift Container Platform debugging information for a hosted cluster by using the multicluster engine Operator web console or by using the CLI.
Gathering data for a hosted cluster by using the CLI¶
You can gather OpenShift Container Platform debugging information for a hosted cluster by using the command-line interface (CLI).
Prerequisites
- You must have
cluster-adminaccess to the management cluster. - You need the
namevalue for theHostedClusterresource and the namespace where the CR is deployed. - You must have the
hcpcommand-line interface installed. For more information, see "Installing the hosted control planes command-line interface". - You must have the OpenShift CLI (
oc) installed. - You must ensure that the
kubeconfigfile is loaded and is pointing to the management cluster.
Procedure
-
Generate the
kubeconfigfile by entering the following command: -
After you save the
kubeconfigfile, you can access the hosted cluster by entering the following example command: -
Collect the must-gather information by entering the following command:
Gathering data for a hosted cluster by using the web console¶
You can gather OpenShift Container Platform debugging information for a hosted cluster by using the multicluster engine Operator web console.
Prerequisites
- You must have
cluster-adminaccess to the management cluster. - You need the
namevalue for theHostedClusterresource and the namespace where the CR is deployed. - You must have the
hcpcommand-line interface installed. For more information, see "Installing the hosted control planes command-line interface". - You must have the OpenShift CLI (
oc) installed. - You must ensure that the
kubeconfigfile is loaded and is pointing to the management cluster.
Procedure
-
In the web console, select All Clusters and select the cluster you want to troubleshoot.
-
In the upper-right corner, select Download kubeconfig.
-
Export the downloaded
kubeconfigfile. -
Collect the must-gather information by entering the following command:
Entering the must-gather command in a disconnected environment¶
When you need to troubleshoot an issue in a disconnected environment, you can gather information by running the must-gather command. The command generates output for the management cluster and the hosted cluster.
Procedure
-
In a disconnected environment, mirror the Red Hat Operator catalog images into their mirror registry. For more information, see "Install on disconnected networks".
-
Run the following command to extract logs that reference the image from their mirror registry:
Additional resources
Troubleshooting hosted clusters on OpenShift Virtualization¶
When you troubleshoot a hosted cluster on OpenShift Virtualization, start with the top-level HostedCluster and NodePool resources and then work down the stack until you find the root cause. The following steps can help you discover the root cause of common issues.
Troubleshooting HostedCluster resource stuck in a partial state¶
If a hosted control plane is not coming fully online because a HostedCluster resource is pending, identify the problem by checking prerequisites, resource conditions, and node and Operator status.
Procedure
-
Ensure that you meet all of the prerequisites for a hosted cluster on OpenShift Virtualization.
-
View the conditions on the
HostedClusterandNodePoolresources for validation errors that prevent progress. -
By using the
kubeconfigfile of the hosted cluster, inspect the status of the hosted cluster:- View the output of the
oc get clusteroperatorscommand to see which cluster Operators are pending. - View the output of the
oc get nodescommand to ensure that worker nodes are ready.
- View the output of the
Identifying why no compute nodes are registered¶
If a hosted control plane is not coming fully online because the hosted control plane has no compute nodes registered, identify the problem by checking the status of various parts of the hosted control plane.
Procedure
-
View the
HostedClusterandNodePoolconditions for failures that indicate what the problem might be. -
Enter the following command to view the KubeVirt compute node virtual machine (VM) status for the
NodePoolresource: -
If the VMs are stuck in the provisioning state, enter the following command to view the CDI import pods within the VM namespace for clues about why the importer pods have not completed:
-
If the VMs are stuck in the starting state, enter the following command to view the status of the virt-launcher pods:
If the virt-launcher pods are in a pending state, investigate why the pods are not being scheduled. For example, not enough resources might exist to run the virt-launcher pods.
-
If the VMs are running but they are not registered as compute nodes, use the web console to gain VNC access to one of the affected VMs. The VNC output indicates whether the ignition configuration was applied. If a VM cannot access the hosted control plane ignition server on startup, the VM cannot be provisioned correctly.
-
If the ignition configuration was applied but the VM is still not registering as a node, see "Identifying the problem: Access the VM console logs" to learn how to access the VM console logs during startup.
Additional resources
Identifying why compute nodes are not ready¶
During cluster creation, nodes enter the NotReady state temporarily while the networking stack is rolled out. This part of the process is normal. However, if this part of the process takes longer than 15 minutes, identify the problem by investigating the node object and pods.
Procedure
-
Enter the following command to view the conditions on the node object and determine why the node is not ready:
-
Enter the following command to look for failing pods within the cluster:
Identifying why ingress and console cluster Operators are not coming online¶
If a hosted control plane is not coming fully online because the Ingress and console cluster Operators are not online, check the wildcard DNS routes and load balancer.
Procedure
-
If the cluster uses the default Ingress behavior, enter the following command to ensure that wildcard DNS routes are enabled on the OpenShift Container Platform cluster that the virtual machines (VMs) are hosted on:
-
If you use a custom base domain for the hosted control plane, complete the following steps:
- Ensure that the load balancer is targeting the VM pods correctly.
- Ensure that the wildcard DNS entry is targeting the load balancer IP address.
Identifying why load balancer services for the hosted cluster are unavailable¶
If a hosted control plane is not coming fully online because the load balancer services are not becoming available, check events, details, and the Kubernetes Cluster Configuration Manager (KCCM) pod.
Procedure
-
Look for events and details that are associated with the load balancer service within the hosted cluster.
-
By default, load balancers for the hosted cluster are handled by the kubevirt-cloud-controller-manager within the hosted control plane namespace. Ensure that the KCCM pod is online and view its logs for errors or warnings. To identify the KCCM pod in the hosted control plane namespace, enter the following command:
Identifying why hosted cluster PVCs are not available¶
If a hosted control plane is not coming fully online because the persistent volume claims (PVCs) for a hosted cluster are not available, check the PVC events and details, and component logs.
Procedure
-
Look for events and details that are associated with the PVC to understand which errors are occurring.
-
If a PVC is failing to attach to a pod, view the logs for the kubevirt-csi-node
daemonsetcomponent within the hosted cluster to further investigate the problem. To identify the kubevirt-csi-node pods for each node, enter the following command: -
If a PVC cannot bind to a persistent volume (PV), view the logs of the kubevirt-csi-controller component within the hosted control plane namespace. To identify the kubevirt-csi-controller pod within the hosted control plane namespace, enter the following command:
Identifying why VM nodes are not joining the cluster¶
If a hosted control plane is not coming fully online because the virtual machine (VM) nodes are not correctly joining the cluster, access the VM console logs.
Procedure
- To access the VM console logs, complete the steps in "How to get serial console logs for VMs part of OpenShift Virtualization Hosted Control Plane clusters".
Additional resources
Resolving RHCOS image mirroring failures¶
For hosted control planes on OpenShift Virtualization in a disconnected environment, if oc-mirror fails to automatically mirror the Red Hat Enterprise Linux CoreOS (RHCOS) image to the internal registry, you can manually mirror the RHCOS image to the internal registry.
When you create your first hosted cluster, the Kubevirt virtual machine does not boot, because the boot image is not available in the internal registry.
To resolve this issue, manually mirror the RHCOS image to the internal registry.
Procedure
-
Get the internal registry name by running the following command:
-
Get a payload image by running the following command:
-
Extract the
0000_50_installer_coreos-bootimages.yamlfile that contains boot images from your payload image on the hosted cluster. Replace<payload_image>with the name of your payload image. Run the following command: -
Get the RHCOS image by running the following command:
-
Mirror the RHCOS image to your internal registry by running the following command:
- Replace
<rhcos_image>with your RHCOS image; for example,quay.io/openshift-release-dev/ocp-v4.0-art-dev@sha256:d9643ead36b1c026be664c9c65c11433c6cdf71bfd93ba229141d134a4a6dd94. - Replace
<internal_registry>with the name of your internal registry; for example,virthost.ostest.test.metalkube.org:5000/localimages/ocp-v4.0-art-dev.
- Replace
-
Create a YAML file named
rhcos-boot-kubevirt.yamlthat defines theImageDigestMirrorSetobject. See the following example configuration:apiVersion: config.openshift.io/v1 kind: ImageDigestMirrorSet metadata: name: rhcos-boot-kubevirt spec: repositoryDigestMirrors: - mirrors: - virthost.ostest.test.metalkube.org:5000/localimages/ocp-v4.0-art-dev source: quay.io/openshift-release-dev/ocp-v4.0-art-devspec.repositoryDigestMirrors.mirrorsspecifies the name of your internal registry.spec.repositoryDigestMirrors.sourcespecifies your RHCOS image without its digest.
-
Apply the
rhcos-boot-kubevirt.yamlfile to create theImageDigestMirrorSetobject by running the following command:
Returning non-bare-metal clusters to the late binding pool¶
If you are using late binding managed clusters without BareMetalHosts, you must complete additional manual steps to delete a late binding cluster and return the nodes back to the Discovery ISO.
For late binding managed clusters without BareMetalHosts, removing cluster information does not automatically return all nodes to the Discovery ISO.
To unbind the non-bare-metal nodes with late binding, complete the following steps.
Procedure
- Remove the cluster information. For more information, see "Removing a cluster from management".
- Clean the root disks.
- Reboot manually with the Discovery ISO.
Additional resources
Troubleshooting hosted clusters on bare metal¶
If you encounter issues with hosted control planes on bare metal, review the troubleshooting procedures to diagnose and resolve them.
Determining why nodes are not added to a hosted cluster on bare metal¶
When you scale up a hosted cluster with nodes that were provisioned by using Assisted Installer, the host fails to pull the ignition with a URL that contains port 22642. That URL is invalid for hosted control planes and indicates that an issue exists with the cluster.
Procedure
-
To determine the issue, review the assisted-service logs by entering the following command:
Replace
<assisted_service_pod_name>with the Assisted Service pod name. -
In the logs, find errors that resemble these examples:
-
To fix this issue, see "Add the pull secret to the namespace" in the multicluster engine for Kubernetes Operator documentation.
Note
To use hosted control planes, you must have multicluster engine Operator installed, either as a standalone Operator or as part of Red Hat Advanced Cluster Management. Because the Operator has a close association with Red Hat Advanced Cluster Management, the documentation for the Operator is published within that product’s documentation. Even if you do not use Red Hat Advanced Cluster Management, the parts of its documentation that cover multicluster engine Operator are relevant to hosted control planes.
Additional resources
Restarting hosted control plane components¶
If you are an administrator for hosted control planes, you can use the hypershift.openshift.io/restart-date annotation to restart all control plane components for a particular HostedCluster resource.
For example, you might need to restart control plane components for certificate rotation.
Procedure
-
To restart a control plane, annotate the
HostedClusterresource by entering the following command:$ oc annotate hostedcluster \ -n <hosted_cluster_namespace> \ <hosted_cluster_name> \ hypershift.openshift.io/restart-date=$(date --iso-8601=seconds)The control plane is restarted whenever the value of the annotation changes. The
datecommand serves as the source of a unique string. The annotation is treated as a string, not a timestamp.
Verification
After you restart a control plane, the following hosted control planes components are typically restarted:
Note
You might see some additional components restarting as a side effect of changes implemented by the other components.
- catalog-operator
- certified-operators-catalog
- cluster-api
- cluster-autoscaler
- cluster-policy-controller
- cluster-version-operator
- community-operators-catalog
- control-plane-operator
- hosted-cluster-config-operator
- ignition-server
- ingress-operator
- konnectivity-agent
- konnectivity-server
- kube-apiserver
- kube-controller-manager
- kube-scheduler
- machine-approver
- oauth-openshift
- olm-operator
- openshift-apiserver
- openshift-controller-manager
- openshift-oauth-apiserver
- packageserver
- redhat-marketplace-catalog
- redhat-operators-catalog
Pausing the reconciliation of a hosted cluster and hosted control plane¶
If you are a cluster instance administrator, you can pause the reconciliation of a hosted cluster and hosted control plane. You might want to pause reconciliation when you back up and restore an etcd database or when you need to debug problems with a hosted cluster or hosted control plane.
Procedure
-
To pause reconciliation for a hosted cluster and hosted control plane, populate the
pausedUntilfield of theHostedClusterresource.-
To pause the reconciliation until a specific time, enter the following command:
$ oc patch -n <hosted_cluster_namespace> \ hostedclusters/<hosted_cluster_name> \ -p '{"spec":{"pausedUntil":"<timestamp>"}}' \ --type=mergeReplace
<timestamp>with a timestamp in the RFC339 format; for example,2024-03-03T03:28:48Z. The reconciliation is paused until the specified time is passed. -
To pause the reconciliation indefinitely, enter the following command:
$ oc patch -n <hosted_cluster_namespace> \ hostedclusters/<hosted_cluster_name> \ -p '{"spec":{"pausedUntil":"true"}}' \ --type=mergeThe reconciliation is paused until you remove the field from the
HostedClusterresource.When the pause reconciliation field is populated for the
HostedClusterresource, the field is automatically added to the associatedHostedControlPlaneresource.
-
-
To remove the
pausedUntilfield, enter the following patch command:
Resolving agent service failures for hosted control planes on IBM Z¶
In some cases, agents might fail to join the cluster after booting the machines with the boot artifacts.
You can confirm this issue by checking the agent.service logs for the following error:
Error: copying system image from manifest list: Source image rejected: A signature was required, but no signature exists
This issue occurs because image signature verification fails when no signature is present. As a workaround, you can disable signature verification by modifying the container policy.
Procedure
-
Add the
ignitionConfigOverridefield in theInfraEnvmanifest to override the/etc/containers/policy.jsonfile. This disables signature verification for container images. -
Replace the base64-encoded content in the
ignitionConfigOverridewith the required/etc/containers/policy.jsonconfiguration according to your image registries. See the following example:Example{ "default": [ { "type": "insecureAcceptAnything" } ], "transports": { "docker": { "<REGISTRY1>": [ { "type": "insecureAcceptAnything" } ], "REGISTRY2": [ { "type": "insecureAcceptAnything" } ] }, "docker-daemon": { "": [ { "type": "insecureAcceptAnything" } ] } } }Example InfraEnv manifest with ignitionConfigOverrideapiVersion: agent-install.openshift.io/v1beta1 kind: InfraEnv metadata: name: <hosted_cluster_name> namespace: <hosted_control_plane_namespace> spec: cpuArchitecture: s390x pullSecretRef: name: pull-secret sshAuthorizedKey: <ssh_public_key> ignitionConfigOverride: '{"ignition":{"version":"3.2.0"},"storage":{"files":[{"path":"/etc/containers/policy.json","mode":420,"overwrite":true,"contents":{"source":"data:text/plain;charset=utf-8;base64,ewogICAgImRlZmF1bHQiOiBbCiAgICAgICAgewogICAgICAgICAgICAidHlwZSI6ICJpbnNlY3VyZUFjY2VwdEFueXRoaW5nIgogICAgICAgIH0KICAgIF0sCiAgICAidHJhbnNwb3J0cyI6CiAgICAgICAgewogICAgICAgICAgICAiZG9ja2VyLWRhZW1vbiI6CiAgICAgICAgICAgICAgICB7CiAgICAgICAgICAgICAgICAgICAgIiI6IFt7InR5cGUiOiJpbnNlY3VyZUFjY2VwdEFueXRoaW5nIn1dCiAgICAgICAgICAgICAgICB9CiAgICAgICAgfQp9"}}]}}'
Known limitations for internal subnets for hosted clusters¶
Several known limitations exist for internal subnets on hosted clusters.
- IPv6 subnets are not supported.
- The hosted control planes command-line interface,
hcp, might not have native flags for the subnet fields. Manual YAML editing oroc patchis required. - Modifying OVN subnets after you create a cluster triggers a rollout of OVN components, which might cause brief network disruptions.
- You cannot modify OVN subnet configuration while a cluster update is in progress or scheduled.
Configuring the ovnKubernetesConfig object fails with an error¶
When you try to configure the ovnKubernetesConfig object on a hosted cluster by using a different network type, such as OpenShiftSDN, an error occurs because hosted control planes works only with the OVNKubernetes network type.
Procedure
-
Verify the network type of your hosted cluster by entering the following command:
Setting CIDR values in internal subnet fields¶
If the internalJoinSubnet field and the internalTransitSwitchSubnet field are set to the same classless inter-domain routing (CIDR) values, an error occurs.
Procedure
-
Use different subnets for each field, as shown in the following example:
Ensuring a valid IPv4 CIDR format¶
If you do not specify subnets in a valid classless inter-domain range (CIDR) format, an error occurs.
Procedure
-
Ensure that the CIDR format follows the following format:
where:
X- is a value from
0to255. The first octet must not be0. Y- is a value from
0to30.
Avoiding an overlap between OVN subnets and CIDR values¶
If the configured OVN subnets overlap with the machine classless inter-domain routing (CIDR), service CIDR, cluster network CIDR, or with each other, an error occurs.
Procedure
-
Use subnets that do not overlap with any network CIDR. You can use a CIDR calculator to verify that no overlaps exist.
Example of configuration with no overlapsspec: networking: machineCIDR: 10.0.0.0/16 serviceCIDR: 172.30.0.0/16 clusterNetwork: - cidr: 10.128.0.0/14 operatorConfiguration: clusterNetworkOperator: ovnKubernetesConfig: ipv4: internalJoinSubnet: "100.99.0.0/16" internalTransitSwitchSubnet: "100.69.0.0/16"
Resolving a stuck OVN rollout¶
After you change an existing configuration, the OVN component rollout might take a long time or encounter issues.
Procedure
-
Check the status of the
ovnkube-nodeDaemonSet rollout by entering the following command: -
Check the pod logs for errors by entering the following command:
If the rollout is stuck, you might need to revert the configuration change.
Troubleshooting connectivity for hosted control planes¶
By using connectivity metrics, you can diagnose whether any issues are related to connectivity from a hosted control plane to a data plane.
Troubleshooting connectivity from the control plane to the data plane¶
To diagnose connectivity issues from a hosted control plane to the compute nodes in a data plane, check the status of the DataPlaneConnectionAvailable condition.
If the status of the DataPlaneConnectionAvailable condition is True, the control plane can successfully reach the data plane nodes through the konnectivity-agent pods. If the status is False, take the following steps to determine why the control plane cannot reach the data plane.
Procedure
- Check the network policies that might block the
Konnectivityservice traffic. - Review the firewall rules between the control plane and the data plane.
- View the status of the
konnectivity-agentpods in the data plane. - In the control plane, review the
Konnectivityserver deployment.
Additional resources