Virtual machine health checks
Define probes and watchdogs in the VirtualMachine resource to configure virtual machine (VM) health checks. Health checks monitor and report the internal state of a VM.
You can configure VM health checks by defining readiness and liveness probes in the VirtualMachine resource.
About readiness and liveness probes
Use readiness and liveness probes to detect and handle unhealthy virtual machines (VMs). You can include one or more probes in the specification of the VM to ensure that traffic does not reach a VM that is not ready for it and that a new VM is created when a VM becomes unresponsive.
A readiness probe determines whether a VM is ready to accept service requests. If the probe fails, the VM is removed from the list of available endpoints until the VM is ready.
A liveness probe determines whether a VM is responsive. If the probe fails, the VM is deleted and a new VM is created to restore responsiveness.
You can configure readiness and liveness probes by setting the spec.readinessProbe and the spec.livenessProbe fields of the VirtualMachine object. These fields support the following tests:
- HTTP GET
- The probe determines the health of the VM by using a web hook. The test is successful if the HTTP response code is between 200 and 399. You can use an HTTP GET test with applications that return HTTP status codes when they are completely initialized.
- TCP socket
- The probe attempts to open a socket to the VM. The VM is only considered healthy if the probe can establish a connection. You can use a TCP socket test with applications that do not start listening until initialization is complete.
- Guest agent ping
- The probe uses the
guest-pingcommand to determine if the QEMU guest agent is running on the virtual machine.
Defining an HTTP readiness probe
You can define an HTTP readiness probe by setting the spec.readinessProbe.httpGet field of the virtual machine (VM) configuration.
Prerequisites
- You have installed the OpenShift CLI (
oc).
Procedure
-
Include details of the readiness probe in the VM configuration file. Sample readiness probe with an HTTP GET test:
apiVersion: kubevirt.io/v1kind: VirtualMachinemetadata:annotations:name: fedora-vmnamespace: example-namespace# ...spec:template:spec:readinessProbe:httpGet:port: 1500path: /healthzhttpHeaders:- name: Custom-Headervalue: AwesomeinitialDelaySeconds: 120periodSeconds: 20timeoutSeconds: 10failureThreshold: 3successThreshold: 3# ...spec.template.spec.readinessProbe.httpGetdefines the HTTP GET request to perform to connect to the VM.spec.template.spec.readinessProbe.httpGet.portdefines the port of the VM that the probe queries. In the above example, the probe queries port 1500.spec.template.spec.readinessProbe.httpGet.pathdefines the path to access on the HTTP server. In the above example, if the handler for the server’s /healthz path returns a success code, the VM is considered to be healthy. If the handler returns a failure code, the VM is removed from the list of available endpoints.spec.template.spec.readinessProbe.initialDelaySecondsdefines the time, in seconds, after the VM starts before the readiness probe is initiated.spec.template.spec.readinessProbe.periodSecondsdefines the delay, in seconds, between performing probes. The default delay is 10 seconds. This value must be greater thantimeoutSeconds.spec.template.spec.readinessProbe.timeoutSecondsdefines the number of seconds of inactivity after which the probe times out and the VM is assumed to have failed. The default value is 1. This value must be lower thanperiodSeconds.spec.template.spec.readinessProbe.failureThresholddefines the number of times that the probe is allowed to fail. The default is 3. After the specified number of attempts, the pod is markedUnready.spec.template.spec.readinessProbe.successThresholddefines the number of times that the probe must report success, after a failure, to be considered successful. The default is 1.
-
Create the VM by running the following command:
$ oc create -f <file_name>.yaml
Defining a TCP readiness probe
You can define a TCP readiness probe by setting the spec.readinessProbe.tcpSocket field of the virtual machine (VM) configuration.
Prerequisites
- You have installed the OpenShift CLI (
oc).
Procedure
-
Include details of the TCP readiness probe in the VM configuration file. Sample readiness probe with a TCP socket test:
apiVersion: kubevirt.io/v1kind: VirtualMachinemetadata:annotations:name: fedora-vmnamespace: example-namespace# ...spec:template:spec:readinessProbe:initialDelaySeconds: 120periodSeconds: 20tcpSocket:port: 1500timeoutSeconds: 10# ...spec.template.spec.readinessProbe.initialDelaySecondsdefines the time, in seconds, after the VM starts before the readiness probe is initiated.spec.template.spec.readinessProbe.periodSecondsdefines the delay, in seconds, between performing probes. The default delay is 10 seconds. This value must be greater thantimeoutSeconds.spec.template.spec.readinessProbe.tcpSocketdefines the TCP action to perform.spec.template.spec.readinessProbe.tcpSocket.portdefines the port of the VM that the probe queries.spec.template.spec.readinessProbe.timeoutSecondsdefines the number of seconds of inactivity after which the probe times out and the VM is assumed to have failed. The default value is 1. This value must be lower thanperiodSeconds.
-
Create the VM by running the following command:
$ oc create -f <file_name>.yaml
Defining an HTTP liveness probe
Define an HTTP liveness probe by setting the spec.livenessProbe.httpGet field of the virtual machine (VM) configuration. You can define both HTTP and TCP tests for liveness probes in the same way as readiness probes. This procedure configures a sample liveness probe with an HTTP GET test.
Prerequisites
- You have installed the OpenShift CLI (
oc).
Procedure
-
Include details of the HTTP liveness probe in the VM configuration file. Sample liveness probe with an HTTP GET test:
apiVersion: kubevirt.io/v1kind: VirtualMachinemetadata:annotations:name: fedora-vmnamespace: example-namespace# ...spec:template:spec:livenessProbe:initialDelaySeconds: 120periodSeconds: 20httpGet:port: 1500path: /healthzhttpHeaders:- name: Custom-Headervalue: AwesometimeoutSeconds: 10# ...spec.tenmplate.spec.livenessProbe.initialDelaySecondsdefines the time, in seconds, after the VM starts before the liveness probe is initiated.spec.tenmplate.spec.livenessProbe.periodSecondsdefines the delay, in seconds, between performing probes. The default delay is 10 seconds. This value must be greater thantimeoutSeconds.spec.tenmplate.spec.livenessProbe.httpGetdefines the HTTP GET request to perform to connect to the VM.spec.tenmplate.spec.livenessProbe.httpGet.portdefines the port of the VM that the probe queries. In the above example, the probe queries port 1500. The VM installs and runs a minimal HTTP server on port 1500 via cloud-init.spec.tenmplate.spec.livenessProbe.httpGet.pathdefines the path to access on the HTTP server. In the above example, if the handler for the server’s/healthzpath returns a success code, the VM is considered to be healthy. If the handler returns a failure code, the VM is deleted and a new VM is created.spec.tenmplate.spec.livenessProbe.timeoutSecondsdefines the number of seconds of inactivity after which the probe times out and the VM is assumed to have failed. The default value is 1. This value must be lower thanperiodSeconds.
-
Create the VM by running the following command:
$ oc create -f <file_name>.yaml
About watchdogs
Watchdog devices monitor guest operating system responsiveness and trigger recovery actions when a virtual machine becomes unresponsive.
A watchdog device continuously monitors the agent running on a virtual machine (VM). When the guest operating system becomes unresponsive, the watchdog can trigger one of three recovery actions depending on how it is configured.
The poweroff action causes the VM to power down immediately. If spec.runStrategy is not set to manual, the VM automatically reboots after powering down.
The reset action reboots the VM in place without allowing the guest operating system to react. When using the reset action, be aware that the reboot time might cause liveness probes to time out. If cluster-level protections detect a failed liveness probe, the VM might be forcibly rescheduled, which increases the overall reboot time.
The shutdown action initiates a graceful shutdown by stopping all services before powering down the VM.
Watchdog functionality is not available for Windows VMs.
To implement watchdog functionality, you must configure the watchdog device for the VM and install the watchdog agent on the guest operating system.
Configuring a watchdog device for the virtual machine
You configure a watchdog device for the virtual machine (VM).
Prerequisites
- For
x86systems, the VM must use a kernel that works with thei6300esbwatchdog device. If you uses390xarchitecture, the kernel must be enabled fordiag288. Red Hat Enterprise Linux (RHEL) images supporti6300esbanddiag288. - You have installed the OpenShift CLI (
oc).
Procedure
-
Create a
YAMLfile with the following contents:apiVersion: kubevirt.io/v1kind: VirtualMachinemetadata:labels:kubevirt.io/vm: <vm-label>name: <vm-name>spec:runStrategy: Haltedtemplate:metadata:labels:kubevirt.io/vm: <vm-label>spec:domain:devices:watchdog:name: <watchdog><watchdog-device-model>:action: "poweroff"# ...-
spec.template.spec.domain.devices.watchdog.name.<watchdog-device-model>defines the watchdog device model to use. Forx86specifyi6300esb. Fors390xspecifydiag288. -
spec.template.spec.domain.devices.watchdog.name.<watchdog-device-model>.actiondefines the watchdog device action. Specifypoweroff,reset, orshutdown. Theshutdownaction requires that the guest virtual machine is responsive to ACPI signals. Usingshutdownis not recommended. The example above configures the watchdog device on a VM with thepoweroffaction and exposes the device as/dev/watchdog.This device can now be used by the watchdog binary.
-
-
Apply the YAML file to your cluster by running the following command:
$ oc apply -f <file_name>.yaml
Verification
-
Run the following command to verify that the VM is connected to the watchdog device:
warningVerification steps are provided for testing watchdog functionality only and must not be run on production machines.
$ lspci | grep watchdog -i -
Run one of the following commands to confirm the watchdog is active:
- Trigger a kernel panic:
# echo c > /proc/sysrq-trigger
- Stop the watchdog service:
# pkill -9 watchdog
- Trigger a kernel panic:
Installing the watchdog agent on the guest
You can install the watchdog agent on the guest and start the watchdog service.
Procedure
- Log in to the virtual machine as root user.
- This step is only required when installing on IBM Z(R) (
s390x). Enablewatchdogby running the following command:# modprobe diag288_wdt - Verify that the
/dev/watchdogfile path is present in the VM by running the following command:# ls /dev/watchdog - Install the
watchdogpackage and its dependencies:# yum install watchdog - Uncomment the following line in the
/etc/watchdog.conffile and save the changes:#watchdog-device = /dev/watchdog - Enable the
watchdogservice to start on boot:# systemctl enable --now watchdog.service
Defining a guest agent ping probe
You can define a guest agent ping probe by setting the spec.readinessProbe.guestAgentPing field of the virtual machine (VM) configuration.
Prerequisites
- The QEMU guest agent must be installed and enabled on the virtual machine.
- You have installed the OpenShift CLI (
oc).
Procedure
-
Include details of the guest agent ping probe in the VM configuration file. For example:
apiVersion: kubevirt.io/v1kind: VirtualMachinemetadata:annotations:name: fedora-vmnamespace: example-namespace# ...spec:template:spec:readinessProbe:guestAgentPing: {}initialDelaySeconds: 120periodSeconds: 20timeoutSeconds: 10failureThreshold: 3successThreshold: 3# ...spec.template.spec.readinessProbe.guestAgentPingdefines the guest agent ping probe to connect to the VM.spec.template.spec.readinessProbe.initialDelaySecondsdefines the time, in seconds, after the VM starts before the guest agent probe is initiated. This value is optional.spec.template.spec.readinessProbe.periodSecondsdefines the delay, in seconds, between performing probes. The default delay is 10 seconds. This value must be greater thantimeoutSeconds. This value is optionalspec.template.spec.readinessProbe.timeoutSecondsdefines the number of seconds of inactivity after which the probe times out and the VM is assumed to have failed. The default value is 1. This value must be lower thanperiodSeconds. This value is optional.spec.template.spec.readinessProbe.failureThresholddefines the number of times that the probe is allowed to fail. The default is 3. After the specified number of attempts, the pod is markedUnready. This value is optional.spec.template.spec.readinessProbe.successThresholddefines the number of times that the probe must report success, after a failure, to be considered successful. The default is 1. This value is optional.
-
Create the VM by running the following command:
$ oc create -f <file_name>.yaml
Additional resources