Configure PCI passthrough
The Peripheral Component Interconnect (PCI) passthrough feature enables you to access and manage hardware devices from a virtual machine (VM). When PCI passthrough is configured, the PCI devices function as if they were physically attached to the guest operating system.
Node preparation for GPU passthrough
You can prevent GPU operands from deploying on worker nodes that you designated for GPU passthrough.
Preventing NVIDIA GPU operands from deploying on nodes
If you use the NVIDIA GPU Operator in your cluster, you can apply the nvidia.com/gpu.deploy.operands=false label to nodes that you do not want to configure for GPU or vGPU operands. This prevents the creation of the pods that configure GPU or vGPU operands and terminates existing pods.
Prerequisites
- The OpenShift CLI (
oc) is installed.
Procedure
-
Label the node by running the following command:
$ oc label node <node_name> nvidia.com/gpu.deploy.operands=falsewhere:
<node_name>- Specifies the name of a node where you do not want to install the NVIDIA GPU operands.
Verification
- Verify that the label was added to the node by running the following command:
$ oc describe node <node_name>
- Optional: If GPU operands were previously deployed on the node, verify their removal.
-
Check the status of the pods in the
nvidia-gpu-operatornamespace by running the following command:$ oc get pods -n nvidia-gpu-operatorExample output:
NAME READY STATUS RESTARTS AGEgpu-operator-59469b8c5c-hw9wj 1/1 Running 0 8dnvidia-sandbox-validator-7hx98 1/1 Running 0 8dnvidia-sandbox-validator-hdb7p 1/1 Running 0 8dnvidia-sandbox-validator-kxwj7 1/1 Terminating 0 9dnvidia-vfio-manager-7w9fs 1/1 Running 0 8dnvidia-vfio-manager-866pz 1/1 Running 0 8dnvidia-vfio-manager-zqtck 1/1 Terminating 0 9d -
Monitor the pod status until the pods with
Terminatingstatus are removed:$ oc get pods -n nvidia-gpu-operatorExample output:
NAME READY STATUS RESTARTS AGEgpu-operator-59469b8c5c-hw9wj 1/1 Running 0 8dnvidia-sandbox-validator-7hx98 1/1 Running 0 8dnvidia-sandbox-validator-hdb7p 1/1 Running 0 8dnvidia-vfio-manager-7w9fs 1/1 Running 0 8dnvidia-vfio-manager-866pz 1/1 Running 0 8d
-
Preparing host devices for PCI passthrough
About preparing a host device for PCI passthrough
To prepare a host device for PCI passthrough by using the CLI, create a MachineConfig object and add kernel arguments to enable the Input-Output Memory Management Unit (IOMMU).
Bind the PCI device to the Virtual Function I/O (VFIO) driver and then expose it in the cluster by editing the permittedHostDevices field of the HyperConverged custom resource (CR). The permittedHostDevices list is empty when you first install the OpenShift Virtualization Operator.
To remove a PCI host device from the cluster by using the CLI, delete the PCI device information from the HyperConverged CR.
Kernel module blocklist for PCI passthrough
For vfio-pci to allocate a PCI device, no other kernel driver can manage that device. If a driver already manages the device, you must add the specific kernel module to a blocklist. Adding a kernel module to a blocklist makes all devices handled by that module unavailable to the host.
Cluster administrators can expose and manage host devices that are permitted to be used in the cluster by using the oc command-line interface (CLI).
You can add a kernel module to the blocklist by creating a MachineConfig object that generates a configuration file in /etc/modprobe.d/ and adds kernel arguments.
The following example shows a MachineConfig object that adds the enic network driver to the blocklist:
apiVersion: machineconfiguration.openshift.io/v1
kind: MachineConfig
metadata:
labels:
machineconfiguration.openshift.io/role: worker
name: 100-blocklist-enic
spec:
config:
ignition:
version: 3.4.0
storage:
files:
- contents:
source: data:,blocklist%20enic%0A
mode: 420
overwrite: true
path: /etc/modprobe.d/blocklist-enic.conf
kernelArguments:
- enic.blocklist=1
- rd.driver.blocklist=enic
Adding kernel arguments to enable the IOMMU driver
You must enable the Input-Output Memory Management Unit (IOMMU) driver before you can configure mediated devices. To enable the IOMMU driver in the kernel, create the MachineConfig object and add the kernel arguments.
Prerequisites
-
You have cluster administrator permissions.
-
Your CPU hardware is Intel or AMD.
noteEnabling IOMMU is not required on
s390xarchitecture. -
You enabled Intel Virtualization Technology for Directed I/O extensions or AMD IOMMU in the BIOS.
-
You have installed the OpenShift CLI (
oc).
Procedure
-
Create a
MachineConfigobject that identifies the kernel argument. The following example shows a kernel argument for an Intel CPU.apiVersion: machineconfiguration.openshift.io/v1kind: MachineConfigmetadata:labels:machineconfiguration.openshift.io/role: workername: 100-worker-iommuspec:config:ignition:version: 3.2.0kernelArguments:- intel_iommu=on# ...metadata.labels.machineconfiguration.openshift.io/rolespecifies that the new kernel argument is applied only to worker nodes.metadata.namespecifies the ranking of this kernel argument (100) among the machine configs and its purpose. If you have an AMD CPU, specify the kernel argument asamd_iommu=on.spec.kernelArgumentsspecifies the kernel argument asintel_iommufor an Intel CPU.
-
Create the new
MachineConfigobject:$ oc create -f 100-worker-kernel-arg-iommu.yaml
Verification
-
Verify that the new
MachineConfigobject was added by entering the following command and observing the output:$ oc get MachineConfigExample output:
NAME IGNITIONVERSION AGE00-master 3.5.0 164m00-worker 3.5.0 164m01-master-container-runtime 3.5.0 164m01-master-kubelet 3.5.0 164m01-worker-container-runtime 3.5.0 164m01-worker-kubelet 3.5.0 164m100-master-chrony-configuration 3.5.0 169m100-master-set-core-user-password 3.5.0 169m100-worker-chrony-configuration 3.5.0 169m100-worker-iommu 3.5.0 14s -
Verify that IOMMU is enabled at the operating system (OS) level by entering the following command:
$ dmesg | grep -i iommu-
If IOMMU is enabled, output is displayed as shown in the following example: Example output:
Intel: [ 0.000000] DMAR: Intel(R) IOMMU DriverAMD: [ 0.000000] AMD-Vi: IOMMU Initialized
-
Binding PCI devices to the VFIO driver
To bind PCI devices to the VFIO (Virtual Function I/O) driver, obtain the values for the vendor ID and the device ID from each device and create a list with the values. Add this list to the MachineConfig object.
The MachineConfig Operator generates the /etc/modprobe.d/vfio.conf on the nodes with the PCI devices, and binds the PCI devices to the VFIO driver.
Prerequisites
-
You added kernel arguments to enable IOMMU for the CPU.
noteEnabling IOMMU is not required on
s390xarchitecture. -
You have installed the OpenShift CLI (
oc).
Procedure
-
Run the
lspcicommand with the name of the GPU accelerator to obtain the vendor ID and the device ID for the PCI device.noteNVIDIA GPU is supported on
x86andaarch64architectures, Intel QAT is supported onx86architecture, and IBM(R) Spyre is supported ons390xarchitecture.$ lspci -nnv | grep -i <gpu_accelerator>Valid values for
<gpu_accelerator>arenvidia,qat, andspyre.Example output:
02:01.0 3D controller [0302]: NVIDIA Corporation GV100GL [Tesla V100 PCIe 32GB] [10de:1eb8] (rev a1) -
Create a Butane config file,
100-worker-vfiopci.bu, binding the PCI device to the VFIO driver.noteThe Butane version you specify in the config file should match the OpenShift Container Platform version and always ends in
0. For example,4.22.0. See "Creating machine configs with Butane" for information about Butane.Example:
variant: openshiftversion: 4.22.0metadata:name: 100-worker-vfiopcilabels:machineconfiguration.openshift.io/role: workerstorage:files:- path: /etc/modprobe.d/vfio.confmode: 0644overwrite: truecontents:inline: |options vfio-pci ids=<vendor_id>:<device_id>- path: /etc/modules-load.d/vfio-pci.confmode: 0644overwrite: truecontents:inline: vfio-pcimetadata.labels.machineconfiguration.openshift.io/role: workerspecifies that the new kernel argument is applied only to compute nodes.storage.files.contents.inline, where the path is/etc/modprobe.d/vfio.conf, specifies the previously determined hexadecimal vendor ID and device ID values to bind a device to the VFIO driver. You can add a list of multiple devices with their vendor and device information.storage.files.path, where thecontents.inlineisvfio-pci, specifies the file that loads thevfio-pcikernel module on the compute nodes.
-
Use Butane to generate a
MachineConfigobject file,100-worker-vfiopci.yaml, containing the configuration to be delivered to the compute nodes:$ butane 100-worker-vfiopci.bu -o 100-worker-vfiopci.yaml -
Apply the
MachineConfigobject to the compute nodes:$ oc apply -f 100-worker-vfiopci.yaml -
Verify that the
MachineConfigobject was added.$ oc get MachineConfigExample output:
NAME GENERATEDBYCONTROLLER IGNITIONVERSION AGE00-master d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h00-worker d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h01-master-container-runtime d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h01-master-kubelet d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h01-worker-container-runtime d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h01-worker-kubelet d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h100-worker-iommu 3.5.0 30s100-worker-vfiopci-configuration 3.5.0 30s
Verification
-
Verify that the VFIO driver is loaded.
$ lspci -nnk -d <vendor_id>:The output confirms that the VFIO driver is being used.
Example output:
04:00.0 3D controller [0302]: NVIDIA Corporation GP102GL [Tesla P40] [10de:1eb8] (rev a1)Subsystem: NVIDIA Corporation Device [10de:1eb8]Kernel driver in use: vfio-pciKernel modules: nouveau
Exposing PCI host devices in the cluster using the CLI
To expose PCI host devices in the cluster, add details about the PCI devices to the spec.permittedHostDevices.pciHostDevices array of the HyperConverged custom resource (CR).
Prerequisites
- You have installed the OpenShift CLI (
oc).
Procedure
-
Edit the
HyperConvergedCR in your default editor by running the following command:$ oc edit hyperconvergeds.v1beta1.hco.kubevirt.io kubevirt-hyperconverged -n openshift-cnv -
Add the PCI device information to the
spec.permittedHostDevices.pciHostDevicesarray. Example configuration file:apiVersion: hco.kubevirt.io/v1beta1kind: HyperConvergedmetadata:name: kubevirt-hyperconvergednamespace: openshift-cnvspec:permittedHostDevices:pciHostDevices:- pciDeviceSelector: "10DE:1DB6"resourceName: "nvidia.com/GV100GL_Tesla_V100"- pciDeviceSelector: "10DE:1EB8"resourceName: "nvidia.com/TU104GL_Tesla_T4"- pciDeviceSelector: "8086:6F54"resourceName: "intel.com/qat"externalResourceProvider: true# ...-
spec.permittedHostDevicesspecifies the host devices that are permitted to be used in the cluster. -
spec.permittedHostDevices.pciHostDevicesspecifies the list of PCI devices available on the node. -
spec.permittedHostDevices.pciHostDevices.pciDeviceSelectorspecifies the vendor ID and the device ID required to identify the PCI device. -
spec.permittedHostDevices.pciHostDevices.resourceNamespecifies the name of a PCI host device. -
spec.permittedHostDevices.pciHostDevices.externalResourceProvideris an optional setting. Setting this field totrueindicates that the resource is provided by an external device plugin. OpenShift Virtualization allows the usage of this device in the cluster but leaves the allocation and monitoring to an external device plugin.noteThe above example snippet shows two PCI host devices that are named
nvidia.com/GV100GL_Tesla_V100andnvidia.com/TU104GL_Tesla_T4added to the list of permitted host devices in theHyperConvergedCR. These devices have been tested and verified to work with OpenShift Virtualization.Example configuration file for an IBM(R) Spyre device on
s390xarchitecture:apiVersion: hco.kubevirt.io/v1beta1kind: HyperConvergedmetadata:name: kubevirt-hyperconvergednamespace: openshift-cnvspec:permittedHostDevices:pciHostDevices:- pciDeviceSelector: "1014:06a8"resourceName: "ibm.com/spyre"# ...
-
-
Save your changes and exit the editor.
Verification
-
Verify that the PCI host devices were added to the node by running the following command. The example output shows that there is one device each associated with the
nvidia.com/GV100GL_Tesla_V100,nvidia.com/TU104GL_Tesla_T4, andintel.com/qatresource names.$ oc describe node <node_name>Example output:
Capacity:cpu: 64devices.kubevirt.io/kvm: 110devices.kubevirt.io/tun: 110devices.kubevirt.io/vhost-net: 110ephemeral-storage: 915128Mihugepages-1Gi: 0hugepages-2Mi: 0memory: 131395264Kinvidia.com/GV100GL_Tesla_V100 1nvidia.com/TU104GL_Tesla_T4 1intel.com/qat: 1pods: 250Allocatable:cpu: 63500mdevices.kubevirt.io/kvm: 110devices.kubevirt.io/tun: 110devices.kubevirt.io/vhost-net: 110ephemeral-storage: 863623130526hugepages-1Gi: 0hugepages-2Mi: 0memory: 130244288Kinvidia.com/GV100GL_Tesla_V100 1nvidia.com/TU104GL_Tesla_T4 1intel.com/qat: 1pods: 250noteWhen using an IBM(R) Spyre device on
s390xarchitecture, the allocated device is shown as follows:ibm.com/spyre: 1.
Removing PCI host devices from the cluster using the CLI
To remove a PCI host device from the cluster, delete the information for that device from the HyperConverged custom resource (CR).
Prerequisites
- You have installed the OpenShift CLI (
oc).
Procedure
-
Edit the
HyperConvergedCR in your default editor by running the following command:$ oc edit hyperconvergeds.v1beta1.hco.kubevirt.io kubevirt-hyperconverged -n openshift-cnv -
Remove the PCI device information from the
spec.permittedHostDevices.pciHostDevicesarray by deleting thepciDeviceSelector,resourceNameandexternalResourceProvider(if applicable), fields for the appropriate device. In this example, the user deletes thenvidia.com/TU104GL_Tesla_T4. Example configuration file:apiVersion: hco.kubevirt.io/v1beta1kind: HyperConvergedmetadata:name: kubevirt-hyperconvergednamespace: openshift-cnvspec:permittedHostDevices:pciHostDevices:- pciDeviceSelector: "10DE:1DB6"resourceName: "nvidia.com/GV100GL_Tesla_V100"# ... -
Save your changes and exit the editor.
Verification
-
Verify that you removed the PCI host device from the node by running the following command. The example output shows that there are zero devices associated with the
nvidia.com/TU104GL_Tesla_T4resource name.$ oc describe node <node_name>Example output:
Capacity:cpu: 64devices.kubevirt.io/kvm: 110devices.kubevirt.io/tun: 110devices.kubevirt.io/vhost-net: 110ephemeral-storage: 915128Mihugepages-1Gi: 0hugepages-2Mi: 0memory: 131395264Kinvidia.com/GV100GL_Tesla_V100 1nvidia.com/TU104GL_Tesla_T4 0pods: 250Allocatable:cpu: 63500mdevices.kubevirt.io/kvm: 110devices.kubevirt.io/tun: 110devices.kubevirt.io/vhost-net: 110ephemeral-storage: 863623130526hugepages-1Gi: 0hugepages-2Mi: 0memory: 130244288Kinvidia.com/GV100GL_Tesla_V100 1nvidia.com/TU104GL_Tesla_T4 0pods: 250
Virtual machine configuration for PCI passthrough
After the PCI devices have been added to the cluster, you can assign them to virtual machines. The PCI devices are now available as if they are physically connected to the virtual machines.
Assigning a PCI device to a virtual machine
When a PCI device is available in a cluster, you can assign it to a virtual machine and enable PCI passthrough.
Procedure
-
Assign the PCI device to a virtual machine as a host device. Example:
apiVersion: kubevirt.io/v1kind: VirtualMachinespec:domain:devices:hostDevices:- deviceName: nvidia.com/TU104GL_Tesla_T4name: hostdevices1spec.template.spec.domain.devices.hostDevices.deviceNamespecifies the name of the PCI device that is permitted on the cluster as a host device. The virtual machine can access this host device. When using an IBM(R) Spyre device ons390xarchitecture, specifyibm.com/spyre:.
Verification
-
Use the following command to verify that the host device is available from the virtual machine.
$ lspci -nnk | grep <gpu_accelerator>Valid values for
<gpu_accelerator>arenvidia,qat, andspyre.Example output:
$ 02:01.0 3D controller [0302]: NVIDIA Corporation GV100GL [Tesla V100 PCIe 32GB] [10de:1eb8] (rev a1)
PCI passthrough on IBM Z
On IBM Z(R) and IBM(R) LinuxONE, you can configure PCI passthrough for Network Express RoCE adapters and IBM(R) Internal Shared Memory (ISM) virtual PCI devices. Both use vfio-pci to pass devices directly to virtual machines.
Configure PCI passthrough for Network Express RoCE adapters on IBM Z
On IBM Z(R) and IBM(R) LinuxONE, you can configure PCI passthrough for Network Express RoCE adapters by using vfio-pci. This procedure applies only when the adapter is configured in RoCE mode through the Hardware Management Console (HMC).
On IBM Z(R) and IBM(R) LinuxONE systems, Network Express adapters can be configured in two modes through the HMC:
- Network Express RoCE: The adapter is exposed as a PCI virtual function, managed by the
mlx5_corekernel driver, and can be passed through to VMs by usingvfio-pci. - Network Express OSA: The adapter uses IBM Z(R) channel-based I/O (OSH PCI functions). OSH functions are not supported on Linux and are not exposed as PCI devices. Do not configure
vfio-pcipassthrough for adapters in OSA mode.
To use vfio-pci for PCI passthrough of RoCE virtual function devices, you must prevent the mlx5_core kernel driver from binding to the device. Because mlx5_core is included in the initramfs image and loads before the root filesystem is mounted, you must add the driver to a blocklist by using both a modprobe.d configuration file and a kernel boot argument. You then bind the device to vfio-pci by using a second MachineConfig.
Prerequisites
- You have installed OpenShift Container Platform 4.21 or later.
- You have installed the OpenShift Virtualization Operator.
- You have cluster administrator permissions.
- You have installed the Butane tool for generating Ignition-compatible
MachineConfigmanifests. - The Network Express adapter is configured in RoCE mode through the HMC.
- You have installed the OpenShift CLI (
oc).
Procedure
-
On each node, identify the RoCE virtual function PCI address and vendor ID by running the following command:
$ lspci | grep -i mellanoxExample output0000:00:00.0 Ethernet controller: Mellanox Technologies ConnectX Family mlx5Gen Virtual Function [15b3:101e]Record the combined PCI vendor and device ID
15b3:101e. This value is used in thevfio-pciandHyperConvergedconfiguration. -
Create a Butane configuration file named
machine-config-roce.buto add themlx5_coredriver to the blocklist:variant: openshiftversion: 4.22.0metadata:name: 100-worker-blocklist-mlx5labels:machineconfiguration.openshift.io/role: masterstorage:files:- path: /etc/modprobe.d/blocklist-mlx5.confmode: 0644overwrite: truecontents:inline: |blocklist mlx5_coreopenshift:kernel_arguments:- rd.driver.blacklist=mlx5_corenoteThe
rd.driver.blacklist=mlx5_corekernel argument is required in addition to themodprobe.dblocklist entry becausemlx5_coreis included in the initramfs image and loads before/etc/modprobe.d/is accessible. The kernel argument blocks the driver at the initramfs stage. -
Convert the Butane file to a
MachineConfigmanifest by running the following command:$ butane machine-config-roce.bu --output machine-config-roce.yaml -
Apply the
MachineConfigto the cluster by running the following command:$ oc apply -f machine-config-roce.yaml -
Watch the
MachineConfigrollout and wait for completion before proceeding:$ oc get mcp master -w -
Create a Butane configuration file named
machine-config-roce1.buto bind RoCE devices tovfio-pci:variant: openshiftversion: 4.22.0metadata:name: 100-worker-vfiopcilabels:machineconfiguration.openshift.io/role: masterstorage:files:- path: /etc/modprobe.d/vfio.confmode: 0644overwrite: truecontents:inline: |options vfio-pci ids=15b3:101e- path: /etc/modules-load.d/vfio-pci.confmode: 0644overwrite: truecontents:inline: vfio-pci -
Convert the Butane file to a
MachineConfigmanifest by running the following command:$ butane machine-config-roce1.bu --output machine-config-roce1.yaml -
Apply the
MachineConfigto the cluster by running the following command:$ oc apply -f machine-config-roce1.yaml -
Watch the
MachineConfigrollout and wait for completion before proceeding:$ oc get mcp master -w -
Edit the
HyperConvergedcustom resource to expose the RoCE device:$ oc edit hyperconverged kubevirt-hyperconverged -n openshift-cnvAdd the device under
spec.virtualization.permittedHostDevices:spec:virtualization:permittedHostDevices:pciHostDevices:- pciDeviceSelector: "15b3:101e"resourceName: ibm.com/roce_vf -
Add a
hostDevicesentry to theVirtualMachinemanifest:apiVersion: kubevirt.io/v1kind: VirtualMachinemetadata:name: <vm_name>namespace: <namespace>spec:template:spec:domain:devices:hostDevices:- deviceName: ibm.com/roce_vfname: hostdevices1 -
Apply the
VirtualMachinemanifest by running the following command:$ oc apply -f <vm_manifest>.yaml
Verification
-
Verify
vfio-pcibinding on the nodes by running the following command:$ lspci -nnk | grep -A 3 "15b3:101e"Example output0000:00:00.0 Ethernet controller [0200]: Mellanox Technologies ConnectX Family mlx5Gen Virtual Function [15b3:101e]Subsystem: Mellanox Technologies Device [15b3:0002]Kernel driver in use: vfio-pciKernel modules: mlx5_coreThe
Kernel driver in use: vfio-pciline confirms that the device is bound tovfio-pcion the host. -
Verify that the RoCE device is present inside the VM by running the following command:
$ lspci -nnk | grep -A 3 "15b3:101e"Example output0001:00:00.0 Ethernet controller [0200]: Mellanox Technologies ConnectX Family mlx5Gen Virtual Function [15b3:101e]Subsystem: Mellanox Technologies Device [15b3:0002]Kernel driver in use: mlx5_coreKernel modules: mlx5_corenoteInside the VM, the RoCE device is managed by the guest kernel’s own
mlx5_coredriver. The host usesvfio-pcito pass the device through. The guest uses its native driver to operate it.
Configure PCI passthrough for IBM ISM virtual PCI devices on IBM Z
On IBM Z(R) and IBM(R) LinuxONE, IBM(R) Internal Shared Memory (ISM) devices are exposed as virtual PCI devices. You can configure PCI passthrough for ISM devices by binding the device to the vfio-pci driver by using a single MachineConfig.
Unlike RoCE passthrough, which requires two separate MachineConfig objects to blocklist mlx5_core and bind vfio-pci, ISM passthrough requires only a single MachineConfig. The ism kernel module does not load during initramfs, so a softdep directive combined with kernel boot arguments is enough to ensure vfio-pci claims the device before ism.
The softdep directive ensures vfio-pci loads before the ism module without completely blocking ism from the host. The ism module remains available in the kernel but does not own the ISM device.
Prerequisites
- You have installed OpenShift Container Platform 4.21 or later.
- You have installed the OpenShift Virtualization Operator.
- You have cluster administrator permissions.
- You have installed the Butane tool for generating Ignition-compatible
MachineConfigmanifests. - The ISM virtual PCI device is present and visible on the PCI bus of the target nodes.
- You have installed the OpenShift CLI (
oc).
Procedure
-
On each node, confirm the ISM device is displayed as a PCI device by running the following command:
$ lspci | grep -i ismExample output0004:00:00.0 Non-VGA unclassified device: IBM Internal Shared Memory (ISM) virtual PCI device -
Retrieve the PCI vendor and device ID by running the following command:
$ lspci -n -s 0004:00:00.0Example output0004:00:00.0 0000: 1014:04edRecord the combined PCI vendor and device ID
1014:04ed. This value is used in thevfio-pciandHyperConvergedconfiguration. -
Create a Butane configuration file named
100-master-vfio-ism.buto bind the ISM device tovfio-pci:variant: openshiftversion: 4.22.0metadata:name: 100-master-vfio-ismlabels:machineconfiguration.openshift.io/role: masterstorage:files:- path: /etc/modprobe.d/vfio-ism.confmode: 0644overwrite: truecontents:inline: |softdep ism pre: vfio-pcioptions vfio-pci ids=1014:04edopenshift:kernel_arguments:- rd.driver.pre=vfio-pci- vfio-pci.ids=1014:04edwhere:
storage.files[].contents.inline softdep ism pre: vfio-pci- Specifies that
vfio-pcimust load before theismmodule. Theismmodule remains available on the host but does not own the device. storage.files[].contents.inline options vfio-pci ids=1014:04ed- Specifies that
vfio-pciclaims devices with this PCI vendor and device ID. openshift.kernel_arguments rd.driver.pre=vfio-pci- Specifies that
vfio-pciloads during initramfs before any other driver. openshift.kernel_arguments vfio-pci.ids=1014:04ed- Specifies the PCI device ID passed directly to
vfio-pciat boot time.
-
Convert the Butane file to a
MachineConfigmanifest by running the following command:$ butane 100-master-vfio-ism.bu -o 100-master-vfio-ism.yaml -
Apply the
MachineConfigto the cluster by running the following command:$ oc apply -f 100-master-vfio-ism.yaml -
Watch the
MachineConfigrollout and wait for completion before proceeding:$ oc get mcp master -wExample output when completeNAME CONFIG UPDATED UPDATING DEGRADED MACHINECOUNT READYMACHINECOUNT UPDATEDMACHINECOUNT DEGRADEDMACHINECOUNT AGEmaster rendered-master-3fb080e65525e49079d4b34e122fb64c True False False 3 3 3 0 27d -
Confirm the ISM device resource is visible and allocatable on the nodes by running the following command:
$ oc describe nodes | grep ibm.com/ismExample outputibm.com/ism: 1ibm.com/ism: 1 -
Edit the
HyperConvergedcustom resource to expose the ISM device by running the following command:$ oc edit hyperconverged kubevirt-hyperconverged -n openshift-cnvAdd the ISM device under
spec.virtualization.permittedHostDevices:spec:virtualization:permittedHostDevices:pciHostDevices:- pciDeviceSelector: "1014:04ed"resourceName: ibm.com/ism -
Verify the HyperConverged Operator accepted the configuration by running the following command:
$ oc get hyperconverged kubevirt-hyperconverged \-n openshift-cnv -o json | jq '.spec.virtualization.permittedHostDevices'Example output{"pciHostDevices": [{"pciDeviceSelector": "1014:04ed","resourceName": "ibm.com/ism"}]} -
Add the ISM device to a
VirtualMachinemanifest:apiVersion: kubevirt.io/v1kind: VirtualMachinemetadata:name: <vm_name>namespace: <namespace>spec:running: truetemplate:spec:domain:devices:hostDevices:- deviceName: ibm.com/ismname: ism-deviceresources:requests:memory: 1Gi -
Apply the
VirtualMachinemanifest by running the following command:$ oc apply -f <vm_manifest>.yaml
Verification
-
Verify
vfio-pcibinding on all control plane nodes by running the following command:$ for node in $(oc get nodes -l node-role.kubernetes.io/master -o name); doecho "=== $node ==="oc debug $node -- chroot /host bash -c \"lspci -nnk | grep -A2 'ISM\|1014:04ed'" 2>/dev/nulldoneExample output=== node/master-0 ===0000:00:00.0 Non-VGA unclassified device [0000]: IBM Internal Shared Memory (ISM) virtual PCI device [1014:04ed]Kernel driver in use: vfio-pciKernel modules: ism=== node/master-1 ===0000:00:00.0 Non-VGA unclassified device [0000]: IBM Internal Shared Memory (ISM) virtual PCI device [1014:04ed]Kernel driver in use: vfio-pciKernel modules: ism=== node/master-2 ===0000:00:00.0 Non-VGA unclassified device [0000]: IBM Internal Shared Memory (ISM) virtual PCI device [1014:04ed]Kernel driver in use: vfio-pciKernel modules: ismKernel modules: ismindicates theismmodule is available in the kernel but is not actively managing the device.vfio-pciowns the device. -
Verify the ISM device is present inside the VM by connecting to the VM console and running the following command:
$ lspci -nnk | grep -i ismExample output0001:00:00.0 Non-VGA unclassified device: IBM Internal Shared Memory (ISM) virtual PCI deviceThe output confirms PCI passthrough. The guest binds the
ismdriver only if the guest image includes that module.
Additional resources