Restarting the cluster gracefully
You can restart your OpenShift Container Platform cluster after a graceful shutdown by powering on nodes and verifying cluster health. The cluster returns to normal operations when nodes and Operators are healthy.
Even though the cluster is expected to be functional after the restart, the cluster might not recover due to unexpected conditions:
- etcd data corruption during shutdown
- Node failure due to hardware
- Network connectivity issues
If your cluster fails to recover, follow the steps in "Restoring to an earlier cluster state".
Restarting the cluster
You can restart the cluster after a graceful shutdown by powering on nodes, uncordoning schedulable nodes, and approving pending certificate signing requests (CSRs) if nodes are not ready. The cluster returns to normal operations after all nodes and Operators are healthy.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - You have gracefully shut down your cluster.
If your cluster fails to recover, follow the steps in "Restoring to an earlier cluster state".
Procedure
-
Turn on the control plane nodes.
- If you are using the
admin.kubeconfigfrom the cluster installation and the API virtual IP address (VIP) is up, complete the following steps:- Set the
KUBECONFIGenvironment variable to theadmin.kubeconfigpath. - Uncordon each control plane node in the cluster by running the following command:
$ oc adm uncordon <node>
- Set the
- If you do not have access to your
admin.kubeconfigcredentials, complete the following steps:- Use SSH to connect to a control plane node.
- Copy the
localhost-recovery.kubeconfigfile to the/rootdirectory. - Use that file to uncordon each control plane node in the cluster by running the following command:
$ oc adm uncordon <node>
- If you are using the
-
Power on any cluster dependencies, such as external storage or a Lightweight Directory Access Protocol (LDAP) server.
-
Start all cluster machines. Use the appropriate method for your cloud environment to start the machines, for example, from the web console for your cloud provider.
Wait approximately 10 minutes before continuing to check the status of control plane nodes.
-
Verify that all control plane nodes are ready by running the following command:
$ oc get nodes -l node-role.kubernetes.io/masterThe control plane nodes are ready if the status is
Ready, as shown in the following output:NAME STATUS ROLES AGE VERSIONip-10-0-168-251.ec2.internal Ready control-plane,master 75m v1.35.0ip-10-0-170-223.ec2.internal Ready control-plane,master 75m v1.35.0ip-10-0-211-16.ec2.internal Ready control-plane,master 75m v1.35.0 -
If the control plane nodes are not ready, then check whether there are any pending CSRs that must be approved.
-
Get the list of current CSRs by running the following command:
$ oc get csr -
Review the details of each CSR to verify that it is valid by running the following command:
$ oc get csr <csr_name> -o jsonpath='{.spec.request}' | base64 -d | openssl req -text -nooutWhen validating the CSR, verify that the following fields match your infrastructure expectations:
-
Subject / Common Name (CN): Must follow the format
system:node:<node_name>, such assystem:node:control-plane-0.example.com. -
Organization (O): Must be exactly
system:nodes. -
Requested Extensions (Extended Key Usage): Must list
TLS Web Client Authentication.Example outputCertificate Request:Data:Version: 1 (0x0)Subject: O = system:nodes, CN = system:node:control-plane-0.example.comSubject Public Key Info:Public Key Algorithm: id-ecPublicKeyPublic-Key: (256 bit)Attributes:Requested Extensions:X509v3 Extended Key Usage:TLS Web Client AuthenticationThe process verifies that the certificate is allowed to be used as a client credential for the node.
-
-
Approve each valid CSR by running the following command:
$ oc adm certificate approve <csr_name>
-
-
After the control plane nodes are ready, verify that all compute nodes are ready by running the following command:
$ oc get nodes -l node-role.kubernetes.io/workerThe compute nodes are ready if the status is
Ready, as shown in the following output:NAME STATUS ROLES AGE VERSIONip-10-0-179-95.ec2.internal Ready worker 64m v1.35.0ip-10-0-182-134.ec2.internal Ready worker 64m v1.35.0ip-10-0-250-100.ec2.internal Ready worker 64m v1.35.0 -
If the compute nodes are not ready, then check whether there are any pending CSRs that must be approved.
-
Get the list of current CSRs by running the following command:
$ oc get csr -
Review the details of each CSR to verify the validity of the CSR by running the following command:
$ oc get csr <csr_name> -o jsonpath='{.spec.request}' | base64 -d | openssl req -text -nooutCompute node CSRs can be for client certificates (kubelet to API) or serving certificates (API to kubelet). Verify that the following fields match your infrastructure expectations:
-
Subject / Common Name (CN): Must follow the format
system:node:<compute_node_name>. -
Organization (O): Must be exactly
system:nodes. -
Extended Key Usage (EKU): Must list
TLS Web Client Authentication(for client requests) orTLS Web Server Authentication(for serving requests). -
Subject Alternative Name (SAN): For serving certificates, this field must contain the correct internal DNS hostname and the internal IP address of the respective compute node.
Example outputCertificate Request:Data:Version: 1 (0x0)Subject: O = system:nodes, CN = system:node:worker-0.example.comSubject Public Key Info:Public Key Algorithm: rsaEncryptionPublic-Key: (2048 bit)Attributes:Requested Extensions:X509v3 Extended Key Usage:TLS Web Server AuthenticationX509v3 Subject Alternative Name:DNS:worker-0.example.com, IP Address:10.0.12.34This process verifies the validity of the certificate as a server credential for cluster communication.
-
-
Approve each valid CSR by running the following command:
$ oc adm certificate approve <csr_name>
-
-
After the control plane and compute nodes are ready, mark all the nodes in the cluster as schedulable by running the following command:
$ for node in $(oc get nodes -o jsonpath='{.items[*].metadata.name}'); do echo ${node} ; oc adm uncordon ${node} ; done
Verification
-
Check that there are no degraded cluster Operators by running the following command:
$ oc get clusteroperatorsExample outputNAME VERSION AVAILABLE PROGRESSING DEGRADED SINCEauthentication 4.22.0 True False False 59mcloud-credential 4.22.0 True False False 85mcluster-autoscaler 4.22.0 True False False 73mconfig-operator 4.22.0 True False False 73mconsole 4.22.0 True False False 62mcsi-snapshot-controller 4.22.0 True False False 66mdns 4.22.0 True False False 76metcd 4.22.0 True False False 76m... -
Check that all nodes are in the
Readystate by running the following command:$ oc get nodesExample outputNAME STATUS ROLES AGE VERSIONip-10-0-168-251.ec2.internal Ready control-plane,master 82m v1.35.0ip-10-0-170-223.ec2.internal Ready control-plane,master 82m v1.35.0ip-10-0-179-95.ec2.internal Ready worker 70m v1.35.0ip-10-0-182-134.ec2.internal Ready worker 70m v1.35.0ip-10-0-211-16.ec2.internal Ready control-plane,master 82m v1.35.0ip-10-0-250-100.ec2.internal Ready worker 69m v1.35.0If the cluster did not start properly, you might need to restore your cluster by using an etcd backup. For more information, see "Restoring to an earlier cluster state".
-
If the CSRs for new compute nodes that are trying to join the cluster are not approved, investigate the problem:
- Check the Machine Approver logs by entering the following command:
$ oc logs -n openshift-cluster-machine-approver -l app=machine-approver
- To list all CSRs, enter the following command:
$ oc get csr -A
- Check the Machine Approver logs by entering the following command:
-
If CSRs are pending, you can manually approve them. Manual approval is needed when certificates expired during extended downtime or after cluster recovery scenarios.
- List any pending CSR requests by entering the following command:
$ oc get csr | grep Pending
- Approve pending CSRs by entering the following command:
$ oc adm certificate approve csrName
- List any pending CSR requests by entering the following command:
-
To troubleshoot other errors related to CSRs, check the Kube Controller Manager logs by entering the following command:
$ oc logs -n openshift-kube-controller-manager -l app=kube-controller-manager
Additional resources