Virtual machine recovery from node failures
To ensure that virtual machines (VMs) recover automatically when a node fails, configure node health checks, automated remediation, and capacity planning. These recommendations come from chaos testing results and help minimize VM downtime during node failure conditions.
How node health monitoring protects virtual machines
By default, virtual machines (VMs) are not automatically recovered to other available nodes when a node fails. To minimize VM downtime and ensure workload resilience, configure node health monitoring and automated remediation.
Without the Node Health Check Operator and a remediation Operator such as Self Node Remediation (SNR) or Fence Agents Remediation (FAR), VMs experience downtime until the node recovers. Apply the following configurations to ensure your workloads are resilient under node failure conditions:
- Configure VMs with
runStrategy: Alwaysto allow recovery to other available nodes. - Deploy and configure the Node Health Check Operator to monitor node status changes and trigger automated remediation.
- Configure a remediation operator, such as SNR or FAR, to recover workloads from failed nodes.
- Plan node capacity to ensure that enough resources are available to host VMs that migrate from failed nodes.
Configuring a VM run strategy by using the CLI
You can configure a run strategy for a virtual machine (VM) by using the command line. The run strategy controls whether a VM automatically restarts after disruptions such as node failures or maintenance events.
Prerequisites
- You have installed the OpenShift CLI (
oc).
Procedure
-
Edit the
VirtualMachineresource by running the following command:$ oc edit vm <vm_name> -n <namespace>Example run strategy:
apiVersion: kubevirt.io/v1kind: VirtualMachinespec:runStrategy: Always# ...
Configuring node health checks for virtual machines
You can configure the Node Health Check Operator to monitor node status and trigger automated remediation when a node becomes unhealthy.
Prerequisites
- You have installed the Node Health Check Operator.
- You have installed a remediation operator.
- You have cluster administrator privileges.
Procedure
-
Create a
NodeHealthCheckcustom resource (CR) to define the health check criteria and remediation strategy:apiVersion: remediation.medik8s.io/v1alpha1kind: NodeHealthCheckmetadata:name: nodehealthcheck-samplespec:minHealthy: <count>pauseRequests:- <pause_test_cluster>remediationTemplate:apiVersion: self-node-remediation.medik8s.io/v1alpha1name: self-node-remediation-automatic-strategy-templatenamespace: openshift-workload-availabilitykind: SelfNodeRemediationTemplateescalatingRemediations:- remediationTemplate:apiVersion: self-node-remediation.medik8s.io/v1alpha1name: self-node-remediation-resource-deletion-templatenamespace: openshift-workload-availabilitykind: SelfNodeRemediationTemplateorder: 1timeout: 300sselector:matchExpressions:- key: node-role.kubernetes.io/workeroperator: ExistsunhealthyConditions:- type: Readystatus: "False"duration: 30s- type: Readystatus: Unknownduration: 30swhere:
-
spec.minHealthydefines the number of worker nodes required to host VMs that migrate from failed nodes.- For critical environments, set this value to the minimum number of nodes that you require to maintain the cluster workload.
- For hyperconverged storage, set this value to one less than your number of nodes to limit remediation to a single worker-node failure and maintain storage stability.
-
spec.remediationTemplatedefines the remediation template to use when the Node Health Check Operator detects an unhealthy node. This example uses the Self Node Remediation Operator. -
spec.escalatingRemediationsdefines escalating remediation strategies. If the initial remediation does not resolve the issue within the specified timeout, the next remediation strategy runs. -
spec.selectordefines the nodes to monitor. This example monitors all worker nodes. -
spec.unhealthyConditionsdefines the parameters to identify an unhealthy node. -
spec.unhealthyConditions.durationdefines the duration that a condition must persist before remediation starts. Set lower values for faster recovery. The following table shows recommended values: RecommendedunhealthyConditionsduration valuesEnvironment Recommended duration Critical 30-60 seconds Standard 60-180 seconds Conservative 180-360 seconds
-
-
Apply the
NodeHealthCheckCR by running the following command:$ oc apply -f nodehealthcheck-sample.yaml
Verification
-
Verify that the
NodeHealthCheckCR exists by running the following command:$ oc get nodehealthcheck nodehealthcheck-sampletipIf remediation does not trigger as expected, verify that the remediation Operator runs correctly, and check the
NodeHealthCheckCR events for errors:$ oc describe nodehealthcheck nodehealthcheck-sample
Node remediation strategies for OpenShift Virtualization
Use the best remediation strategy for your environment to ensure that virtual machine (VM) workloads recover automatically when nodes become unhealthy.
Self Node Remediation (SNR) and Fence Agents Remediation (FAR) are the recommended remediation Operators for OpenShift Virtualization environments. Recovery time depends on the number of VMs on the failed nodes.
- Self Node Remediation (SNR)
- Use for nodes that do not have a management interface, or when the management interface might be unreachable. SNR does not require a management interface to function. Do not use SNR for hyperconverged storage because API unavailability can cause remediation actions that negatively impact hyperconverged storage protection domains.
- Fence Agents Remediation (FAR)
- The fastest remediation Operator for workload recovery time. Use FAR when a baseboard management controller (BMC) interface is available. FAR requires a management interface that is reachable from the Kubernetes pod network.
You must use FAR for hyperconverged storage configurations. For hyperconverged storage, set the Node Health Check
minHealthyproperty to your total node count minus one. With this setting, FAR does not remediate more than a single worker-node failure within the same zone, storage placement, and replication group to maintain storage stability.
- Remediation actions start after the node conditions monitored by the Node Health Check Operator exceed the time specified in
spec.unhealthyConditions[].duration. - Disruption to the node running the
self-node-remediation-controller-managerpod increases recovery times.
Self Node Remediation configuration parameters
You can tune the Self Node Remediation (SNR) Operator configuration to optimize recovery times and ensure availability of VM workloads.
- SelfNodeRemediationTemplate strategies
- The
SelfNodeRemediationTemplatecustom resource (CR) defines the remediation strategy to use when the Node Health Check Operator identifies a node as unhealthy. Configure the remediation strategy in aSelfNodeRemediationTemplateCR:where:apiVersion: self-node-remediation.medik8s.io/v1alpha1kind: SelfNodeRemediationTemplatemetadata:name: self-node-remediation-resource-deletion-templatenamespace: openshift-operatorsspec:template:spec:remediationStrategy: ResourceDeletionspec.remediationStrategydefines the remediation strategy. UseResourceDeletionto delete pods and associated volume attachments on the unhealthy node. UseOutOfServiceTaintto place anout-of-servicetaint on the unhealthy node and delete pods and associated volume attachments. This strategy is faster thanResourceDeletion.
- SelfNodeRemediationConfig parameters
- The
SelfNodeRemediationConfigCR defines the timing and connectivity parameters for the SNR Operator. Tuning these parameters based on your environment conditions helps ensure availability of workloads and prevents unnecessary remediation actions. Configure the timing and connectivity parameters in aSelfNodeRemediationConfigCR:where:apiVersion: self-node-remediation.medik8s.io/v1alpha1kind: SelfNodeRemediationConfigmetadata:name: self-node-remediation-confignamespace: openshift-operatorsspec:safeTimeToAssumeNodeRebootedSeconds: 180watchdogFilePath: /dev/watchdogisSoftwareRebootEnabled: trueapiServerTimeout: 15sapiCheckInterval: 5smaxApiErrorThreshold: 3peerApiServerTimeout: 5speerDialTimeout: 5speerRequestTimeout: 5speerUpdateInterval: 15mspec.safeTimeToAssumeNodeRebootedSecondsdefines the time in seconds to wait before assuming the node has rebooted. Set this value to the node reboot time plus 20% for faster recovery, or plus 40% for a more conservative approach.
Capacity planning for VM failover
Plan node capacity to ensure your cluster has enough resources to host VMs that migrate from failed nodes.
When a node fails and you have enabled remediation, VMs automatically migrate to other available nodes in the cluster. Monitor the following resources on your cluster to ensure adequate capacity for VM failover:
- VM count per node
- Available CPU
- Available memory
- Available disk
- Available network bandwidth
Check the current resource use of your nodes:
$ oc adm top nodes
The number of VMs that a node can host depends on the maxPods value set in the kubelet configuration:
kubeletConfig:
maxPods: 250
Additional resources