Troubleshooting CRI-O container runtime issues¶
Use the following sections to troubleshoot CRI-O container runtime issues.
About CRI-O container runtime engine¶
CRI-O is a Kubernetes-native container engine implementation that integrates closely with the operating system to deliver an efficient and optimized Kubernetes experience. The CRI-O container engine runs as a systemd service on each OpenShift Container Platform cluster node.
When container runtime issues occur, verify the status of the crio systemd service on each node. Gather CRI-O journald unit logs from nodes that have container runtime issues.
Verifying CRI-O runtime engine status¶
You can verify CRI-O container runtime engine status on each cluster node.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - You have installed the OpenShift CLI (
oc).
Procedure
-
Review CRI-O status by querying the
criosystemd service on a node, within a debug pod.-
Start a debug pod for a node:
-
Set
/hostas the root directory within the debug shell. The debug pod mounts the host’s root file system in/hostwithin the pod. By changing the root directory to/host, you can run binaries contained in the host’s executable paths:Note
OpenShift Container Platform 4.22 cluster nodes running Red Hat Enterprise Linux CoreOS (RHCOS) are immutable and rely on Operators to apply cluster changes. Accessing cluster nodes by using SSH is not recommended. However, if the OpenShift Container Platform API is not available, or the kubelet is not properly functioning on the target node,
ocoperations will be impacted. In such situations, it is possible to access nodes usingssh core@<node>.<cluster_name>.<base_domain>instead. -
Check whether the
criosystemd service is active on the node: -
Output a more detailed
crio.servicestatus summary:
-
Gathering CRI-O journald unit logs¶
If you experience CRI-O issues, you can obtain CRI-O journald unit logs from a node.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - Your API service is still functional.
- You have installed the OpenShift CLI (
oc). - You have the fully qualified domain names of the control plane or control plane machines.
Procedure
-
Gather CRI-O journald unit logs. The following example collects logs from all control plane nodes (within the cluster:
-
Gather CRI-O journald unit logs from a specific node:
-
If the API is not functional, review the logs using SSH instead. Replace
<node>.<cluster_name>.<base_domain>with appropriate values:Note
OpenShift Container Platform 4.22 cluster nodes running Red Hat Enterprise Linux CoreOS (RHCOS) are immutable and rely on Operators to apply cluster changes. Accessing cluster nodes by using SSH is not recommended. Before attempting to collect diagnostic data over SSH, review whether the data collected by running
oc adm must gatherand otheroccommands is sufficient instead. However, if the OpenShift Container Platform API is not available, or the kubelet is not properly functioning on the target node,ocoperations will be impacted. In such situations, it is possible to access nodes usingssh core@<node>.<cluster_name>.<base_domain>.
Cleaning CRI-O storage¶
You can manually clear the CRI-O ephemeral storage if you experience the following issues:
-
A node cannot run any pods and this error appears:
-
You cannot create a new container on a working node and the “can’t stat lower layer” error appears:
-
Your node is in the
NotReadystate after a cluster upgrade or if you attempt to reboot it. -
The container runtime implementation (
crio) is not working properly. -
You are unable to start a debug shell on the node using
oc debug node/<node_name>because the container runtime instance (crio) is not working.
Follow this process to completely wipe the CRI-O storage and resolve the errors.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - You have installed the OpenShift CLI (
oc).
Procedure
-
Use
cordonon the node. This is to avoid any workload getting scheduled if the node gets into theReadystatus. You will know that scheduling is disabled whenSchedulingDisabledis in your Status section: -
Drain the node as the cluster-admin user:
Note
The
terminationGracePeriodSecondsattribute of a pod or pod template controls the graceful termination period. This attribute defaults at 30 seconds, but can be customized for each application as necessary. If set to more than 90 seconds, the pod might be marked asSIGKILLedand fail to terminate successfully. -
When the node returns, connect back to the node via SSH or Console. Then connect to the root user:
-
Manually stop the kubelet:
-
Stop the containers and pods:
-
Use the following command to stop the pods that are not in the
HostNetwork. They must be removed first because their removal relies on the networking plugin pods, which are in theHostNetwork. -
Stop all other pods:
-
-
Manually stop the crio services:
-
After you run those commands, you can completely wipe the ephemeral storage:
-
Start the crio and kubelet service:
-
You will know if the clean up worked if the crio and kubelet services are started, and the node is in the
Readystatus: -
Mark the node schedulable. You will know that the scheduling is enabled when
SchedulingDisabledis no longer in status: