Configure PCI passthrough¶
The Peripheral Component Interconnect (PCI) passthrough feature enables you to access and manage hardware devices from a virtual machine (VM). When PCI passthrough is configured, the PCI devices function as if they were physically attached to the guest operating system.
Node preparation for GPU passthrough¶
You can prevent GPU operands from deploying on worker nodes that you designated for GPU passthrough.
Preventing NVIDIA GPU operands from deploying on nodes¶
If you use the NVIDIA GPU Operator in your cluster, you can apply the nvidia.com/gpu.deploy.operands=false label to nodes that you do not want to configure for GPU or vGPU operands. This prevents the creation of the pods that configure GPU or vGPU operands and terminates existing pods.
Prerequisites
- The OpenShift CLI (
oc) is installed.
Procedure
-
Label the node by running the following command:
where:
<node_name>- Specifies the name of a node where you do not want to install the NVIDIA GPU operands.
Verification
-
Verify that the label was added to the node by running the following command:
-
Optional: If GPU operands were previously deployed on the node, verify their removal.
-
Check the status of the pods in the
nvidia-gpu-operatornamespace by running the following command:Example output:
NAME READY STATUS RESTARTS AGE gpu-operator-59469b8c5c-hw9wj 1/1 Running 0 8d nvidia-sandbox-validator-7hx98 1/1 Running 0 8d nvidia-sandbox-validator-hdb7p 1/1 Running 0 8d nvidia-sandbox-validator-kxwj7 1/1 Terminating 0 9d nvidia-vfio-manager-7w9fs 1/1 Running 0 8d nvidia-vfio-manager-866pz 1/1 Running 0 8d nvidia-vfio-manager-zqtck 1/1 Terminating 0 9d -
Monitor the pod status until the pods with
Terminatingstatus are removed:Example output:
-
Preparing host devices for PCI passthrough¶
About preparing a host device for PCI passthrough¶
To prepare a host device for PCI passthrough by using the CLI, create a MachineConfig object and add kernel arguments to enable the Input-Output Memory Management Unit (IOMMU).
Bind the PCI device to the Virtual Function I/O (VFIO) driver and then expose it in the cluster by editing the permittedHostDevices field of the HyperConverged custom resource (CR). The permittedHostDevices list is empty when you first install the OpenShift Virtualization Operator.
To remove a PCI host device from the cluster by using the CLI, delete the PCI device information from the HyperConverged CR.
Kernel module blocklist for PCI passthrough¶
For vfio-pci to allocate a PCI device, no other kernel driver can manage that device. If a driver already manages the device, you must add the specific kernel module to a blocklist. Adding a kernel module to a blocklist makes all devices handled by that module unavailable to the host.
Cluster administrators can expose and manage host devices that are permitted to be used in the cluster by using the oc command-line interface (CLI).
You can add a kernel module to the blocklist by creating a MachineConfig object that generates a configuration file in /etc/modprobe.d/ and adds kernel arguments.
The following example shows a MachineConfig object that adds the enic network driver to the blocklist:
apiVersion: machineconfiguration.openshift.io/v1
kind: MachineConfig
metadata:
labels:
machineconfiguration.openshift.io/role: worker
name: 100-blocklist-enic
spec:
config:
ignition:
version: 3.4.0
storage:
files:
- contents:
source: data:,blocklist%20enic%0A
mode: 420
overwrite: true
path: /etc/modprobe.d/blocklist-enic.conf
kernelArguments:
- enic.blocklist=1
- rd.driver.blocklist=enic
Adding kernel arguments to enable the IOMMU driver¶
You must enable the Input-Output Memory Management Unit (IOMMU) driver before you can configure mediated devices. To enable the IOMMU driver in the kernel, create the MachineConfig object and add the kernel arguments.
Prerequisites
-
You have cluster administrator permissions.
-
Your CPU hardware is Intel or AMD.
Note
Enabling IOMMU is not required on
s390xarchitecture. -
You enabled Intel Virtualization Technology for Directed I/O extensions or AMD IOMMU in the BIOS.
-
You have installed the OpenShift CLI (
oc).
Procedure
-
Create a
MachineConfigobject that identifies the kernel argument. The following example shows a kernel argument for an Intel CPU.apiVersion: machineconfiguration.openshift.io/v1 kind: MachineConfig metadata: labels: machineconfiguration.openshift.io/role: worker name: 100-worker-iommu spec: config: ignition: version: 3.2.0 kernelArguments: - intel_iommu=on # ...metadata.labels.machineconfiguration.openshift.io/rolespecifies that the new kernel argument is applied only to worker nodes.metadata.namespecifies the ranking of this kernel argument (100) among the machine configs and its purpose. If you have an AMD CPU, specify the kernel argument asamd_iommu=on.spec.kernelArgumentsspecifies the kernel argument asintel_iommufor an Intel CPU.
-
Create the new
MachineConfigobject:
Verification
-
Verify that the new
MachineConfigobject was added by entering the following command and observing the output:Example output:
NAME IGNITIONVERSION AGE 00-master 3.5.0 164m 00-worker 3.5.0 164m 01-master-container-runtime 3.5.0 164m 01-master-kubelet 3.5.0 164m 01-worker-container-runtime 3.5.0 164m 01-worker-kubelet 3.5.0 164m 100-master-chrony-configuration 3.5.0 169m 100-master-set-core-user-password 3.5.0 169m 100-worker-chrony-configuration 3.5.0 169m 100-worker-iommu 3.5.0 14s -
Verify that IOMMU is enabled at the operating system (OS) level by entering the following command:
-
If IOMMU is enabled, output is displayed as shown in the following example:
Example output:
-
Binding PCI devices to the VFIO driver¶
To bind PCI devices to the VFIO (Virtual Function I/O) driver, obtain the values for the vendor ID and the device ID from each device and create a list with the values. Add this list to the MachineConfig object.
The MachineConfig Operator generates the /etc/modprobe.d/vfio.conf on the nodes with the PCI devices, and binds the PCI devices to the VFIO driver.
Prerequisites
-
You added kernel arguments to enable IOMMU for the CPU.
Note
Enabling IOMMU is not required on
s390xarchitecture. -
You have installed the OpenShift CLI (
oc).
Procedure
-
Run the
lspcicommand with the name of the GPU accelerator to obtain the vendor ID and the device ID for the PCI device.Note
NVIDIA GPU is supported on
x86andaarch64architectures, Intel QAT is supported onx86architecture, and IBM(R) Spyre is supported ons390xarchitecture.Valid values for
<gpu_accelerator>arenvidia,qat, andspyre.Example output:
-
Create a Butane config file,
100-worker-vfiopci.bu, binding the PCI device to the VFIO driver.Note
The Butane version you specify in the config file should match the OpenShift Container Platform version and always ends in
0. For example,4.22.0. See "Creating machine configs with Butane" for information about Butane.Example:
variant: openshift version: 4.22.0 metadata: name: 100-worker-vfiopci labels: machineconfiguration.openshift.io/role: worker storage: files: - path: /etc/modprobe.d/vfio.conf mode: 0644 overwrite: true contents: inline: | options vfio-pci ids=<vendor_id>:<device_id> - path: /etc/modules-load.d/vfio-pci.conf mode: 0644 overwrite: true contents: inline: vfio-pcimetadata.labels.machineconfiguration.openshift.io/role: workerspecifies that the new kernel argument is applied only to compute nodes.storage.files.contents.inline, where the path is/etc/modprobe.d/vfio.conf, specifies the previously determined hexadecimal vendor ID and device ID values to bind a device to the VFIO driver. You can add a list of multiple devices with their vendor and device information.storage.files.path, where thecontents.inlineisvfio-pci, specifies the file that loads thevfio-pcikernel module on the compute nodes.
-
Use Butane to generate a
MachineConfigobject file,100-worker-vfiopci.yaml, containing the configuration to be delivered to the compute nodes: -
Apply the
MachineConfigobject to the compute nodes: -
Verify that the
MachineConfigobject was added.Example output:
NAME GENERATEDBYCONTROLLER IGNITIONVERSION AGE 00-master d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h 00-worker d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h 01-master-container-runtime d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h 01-master-kubelet d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h 01-worker-container-runtime d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h 01-worker-kubelet d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h 100-worker-iommu 3.5.0 30s 100-worker-vfiopci-configuration 3.5.0 30s
Verification
-
Verify that the VFIO driver is loaded.
The output confirms that the VFIO driver is being used.
Example output:
Exposing PCI host devices in the cluster using the CLI¶
To expose PCI host devices in the cluster, add details about the PCI devices to the spec.permittedHostDevices.pciHostDevices array of the HyperConverged custom resource (CR).
Prerequisites
- You have installed the OpenShift CLI (
oc).
Procedure
-
Edit the
HyperConvergedCR in your default editor by running the following command: -
Add the PCI device information to the
spec.permittedHostDevices.pciHostDevicesarray.Example configuration file:
apiVersion: hco.kubevirt.io/v1beta1 kind: HyperConverged metadata: name: kubevirt-hyperconverged namespace: openshift-cnv spec: permittedHostDevices: pciHostDevices: - pciDeviceSelector: "10DE:1DB6" resourceName: "nvidia.com/GV100GL_Tesla_V100" - pciDeviceSelector: "10DE:1EB8" resourceName: "nvidia.com/TU104GL_Tesla_T4" - pciDeviceSelector: "8086:6F54" resourceName: "intel.com/qat" externalResourceProvider: true # ...-
spec.permittedHostDevicesspecifies the host devices that are permitted to be used in the cluster. -
spec.permittedHostDevices.pciHostDevicesspecifies the list of PCI devices available on the node. -
spec.permittedHostDevices.pciHostDevices.pciDeviceSelectorspecifies the vendor ID and the device ID required to identify the PCI device. -
spec.permittedHostDevices.pciHostDevices.resourceNamespecifies the name of a PCI host device. -
spec.permittedHostDevices.pciHostDevices.externalResourceProvideris an optional setting. Setting this field totrueindicates that the resource is provided by an external device plugin. OpenShift Virtualization allows the usage of this device in the cluster but leaves the allocation and monitoring to an external device plugin.Note
The above example snippet shows two PCI host devices that are named
nvidia.com/GV100GL_Tesla_V100andnvidia.com/TU104GL_Tesla_T4added to the list of permitted host devices in theHyperConvergedCR. These devices have been tested and verified to work with OpenShift Virtualization.Example configuration file for an IBM(R) Spyre device on
s390xarchitecture:
-
-
Save your changes and exit the editor.
Verification
-
Verify that the PCI host devices were added to the node by running the following command. The example output shows that there is one device each associated with the
nvidia.com/GV100GL_Tesla_V100,nvidia.com/TU104GL_Tesla_T4, andintel.com/qatresource names.Example output:
Capacity: cpu: 64 devices.kubevirt.io/kvm: 110 devices.kubevirt.io/tun: 110 devices.kubevirt.io/vhost-net: 110 ephemeral-storage: 915128Mi hugepages-1Gi: 0 hugepages-2Mi: 0 memory: 131395264Ki nvidia.com/GV100GL_Tesla_V100 1 nvidia.com/TU104GL_Tesla_T4 1 intel.com/qat: 1 pods: 250 Allocatable: cpu: 63500m devices.kubevirt.io/kvm: 110 devices.kubevirt.io/tun: 110 devices.kubevirt.io/vhost-net: 110 ephemeral-storage: 863623130526 hugepages-1Gi: 0 hugepages-2Mi: 0 memory: 130244288Ki nvidia.com/GV100GL_Tesla_V100 1 nvidia.com/TU104GL_Tesla_T4 1 intel.com/qat: 1 pods: 250Note
When using an IBM(R) Spyre device on
s390xarchitecture, the allocated device is shown as follows:ibm.com/spyre: 1.
Removing PCI host devices from the cluster using the CLI¶
To remove a PCI host device from the cluster, delete the information for that device from the HyperConverged custom resource (CR).
Prerequisites
- You have installed the OpenShift CLI (
oc).
Procedure
-
Edit the
HyperConvergedCR in your default editor by running the following command: -
Remove the PCI device information from the
spec.permittedHostDevices.pciHostDevicesarray by deleting thepciDeviceSelector,resourceNameandexternalResourceProvider(if applicable), fields for the appropriate device. In this example, the user deletes thenvidia.com/TU104GL_Tesla_T4.Example configuration file:
-
Save your changes and exit the editor.
Verification
-
Verify that you removed the PCI host device from the node by running the following command. The example output shows that there are zero devices associated with the
nvidia.com/TU104GL_Tesla_T4resource name.Example output:
Capacity: cpu: 64 devices.kubevirt.io/kvm: 110 devices.kubevirt.io/tun: 110 devices.kubevirt.io/vhost-net: 110 ephemeral-storage: 915128Mi hugepages-1Gi: 0 hugepages-2Mi: 0 memory: 131395264Ki nvidia.com/GV100GL_Tesla_V100 1 nvidia.com/TU104GL_Tesla_T4 0 pods: 250 Allocatable: cpu: 63500m devices.kubevirt.io/kvm: 110 devices.kubevirt.io/tun: 110 devices.kubevirt.io/vhost-net: 110 ephemeral-storage: 863623130526 hugepages-1Gi: 0 hugepages-2Mi: 0 memory: 130244288Ki nvidia.com/GV100GL_Tesla_V100 1 nvidia.com/TU104GL_Tesla_T4 0 pods: 250
Virtual machine configuration for PCI passthrough¶
After the PCI devices have been added to the cluster, you can assign them to virtual machines. The PCI devices are now available as if they are physically connected to the virtual machines.
Assigning a PCI device to a virtual machine¶
When a PCI device is available in a cluster, you can assign it to a virtual machine and enable PCI passthrough.
Procedure
-
Assign the PCI device to a virtual machine as a host device.
Example:
apiVersion: kubevirt.io/v1 kind: VirtualMachine spec: domain: devices: hostDevices: - deviceName: nvidia.com/TU104GL_Tesla_T4 name: hostdevices1spec.template.spec.domain.devices.hostDevices.deviceNamespecifies the name of the PCI device that is permitted on the cluster as a host device. The virtual machine can access this host device. When using an IBM(R) Spyre device ons390xarchitecture, specifyibm.com/spyre:.
Verification
-
Use the following command to verify that the host device is available from the virtual machine.
Valid values for
<gpu_accelerator>arenvidia,qat, andspyre.Example output:
PCI passthrough on IBM Z¶
On IBM Z(R) and IBM(R) LinuxONE, you can configure PCI passthrough for Network Express RoCE adapters and IBM(R) Internal Shared Memory (ISM) virtual PCI devices. Both use vfio-pci to pass devices directly to virtual machines.
Configure PCI passthrough for Network Express RoCE adapters on IBM Z¶
On IBM Z(R) and IBM(R) LinuxONE, you can configure PCI passthrough for Network Express RoCE adapters by using vfio-pci. This procedure applies only when the adapter is configured in RoCE mode through the Hardware Management Console (HMC).
On IBM Z(R) and IBM(R) LinuxONE systems, Network Express adapters can be configured in two modes through the HMC:
- Network Express RoCE: The adapter is exposed as a PCI virtual function, managed by the
mlx5_corekernel driver, and can be passed through to VMs by usingvfio-pci. - Network Express OSA: The adapter uses IBM Z(R) channel-based I/O (OSH PCI functions). OSH functions are not supported on Linux and are not exposed as PCI devices. Do not configure
vfio-pcipassthrough for adapters in OSA mode.
To use vfio-pci for PCI passthrough of RoCE virtual function devices, you must prevent the mlx5_core kernel driver from binding to the device. Because mlx5_core is included in the initramfs image and loads before the root filesystem is mounted, you must add the driver to a blocklist by using both a modprobe.d configuration file and a kernel boot argument. You then bind the device to vfio-pci by using a second MachineConfig.
Prerequisites
- You have installed OpenShift Container Platform 4.21 or later.
- You have installed the OpenShift Virtualization Operator.
- You have cluster administrator permissions.
- You have installed the Butane tool for generating Ignition-compatible
MachineConfigmanifests. - The Network Express adapter is configured in RoCE mode through the HMC.
- You have installed the OpenShift CLI (
oc).
Procedure
-
On each node, identify the RoCE virtual function PCI address and vendor ID by running the following command:
Example output0000:00:00.0 Ethernet controller: Mellanox Technologies ConnectX Family mlx5Gen Virtual Function [15b3:101e]Record the combined PCI vendor and device ID
15b3:101e. This value is used in thevfio-pciandHyperConvergedconfiguration. -
Create a Butane configuration file named
machine-config-roce.buto add themlx5_coredriver to the blocklist:variant: openshift version: 4.22.0 metadata: name: 100-worker-blocklist-mlx5 labels: machineconfiguration.openshift.io/role: master storage: files: - path: /etc/modprobe.d/blocklist-mlx5.conf mode: 0644 overwrite: true contents: inline: | blocklist mlx5_core openshift: kernel_arguments: - rd.driver.blacklist=mlx5_coreNote
The
rd.driver.blacklist=mlx5_corekernel argument is required in addition to themodprobe.dblocklist entry becausemlx5_coreis included in the initramfs image and loads before/etc/modprobe.d/is accessible. The kernel argument blocks the driver at the initramfs stage. -
Convert the Butane file to a
MachineConfigmanifest by running the following command: -
Apply the
MachineConfigto the cluster by running the following command: -
Watch the
MachineConfigrollout and wait for completion before proceeding: -
Create a Butane configuration file named
machine-config-roce1.buto bind RoCE devices tovfio-pci:variant: openshift version: 4.22.0 metadata: name: 100-worker-vfiopci labels: machineconfiguration.openshift.io/role: master storage: files: - path: /etc/modprobe.d/vfio.conf mode: 0644 overwrite: true contents: inline: | options vfio-pci ids=15b3:101e - path: /etc/modules-load.d/vfio-pci.conf mode: 0644 overwrite: true contents: inline: vfio-pci -
Convert the Butane file to a
MachineConfigmanifest by running the following command: -
Apply the
MachineConfigto the cluster by running the following command: -
Watch the
MachineConfigrollout and wait for completion before proceeding: -
Edit the
HyperConvergedcustom resource to expose the RoCE device:Add the device under
spec.virtualization.permittedHostDevices: -
Add a
hostDevicesentry to theVirtualMachinemanifest: -
Apply the
VirtualMachinemanifest by running the following command:
Verification
-
Verify
vfio-pcibinding on the nodes by running the following command:Example output0000:00:00.0 Ethernet controller [0200]: Mellanox Technologies ConnectX Family mlx5Gen Virtual Function [15b3:101e] Subsystem: Mellanox Technologies Device [15b3:0002] Kernel driver in use: vfio-pci Kernel modules: mlx5_coreThe
Kernel driver in use: vfio-pciline confirms that the device is bound tovfio-pcion the host. -
Verify that the RoCE device is present inside the VM by running the following command:
Example output0001:00:00.0 Ethernet controller [0200]: Mellanox Technologies ConnectX Family mlx5Gen Virtual Function [15b3:101e] Subsystem: Mellanox Technologies Device [15b3:0002] Kernel driver in use: mlx5_core Kernel modules: mlx5_coreNote
Inside the VM, the RoCE device is managed by the guest kernel’s own
mlx5_coredriver. The host usesvfio-pcito pass the device through. The guest uses its native driver to operate it.
Configure PCI passthrough for IBM ISM virtual PCI devices on IBM Z¶
On IBM Z(R) and IBM(R) LinuxONE, IBM(R) Internal Shared Memory (ISM) devices are exposed as virtual PCI devices. You can configure PCI passthrough for ISM devices by binding the device to the vfio-pci driver by using a single MachineConfig.
Unlike RoCE passthrough, which requires two separate MachineConfig objects to blocklist mlx5_core and bind vfio-pci, ISM passthrough requires only a single MachineConfig. The ism kernel module does not load during initramfs, so a softdep directive combined with kernel boot arguments is enough to ensure vfio-pci claims the device before ism.
Note
The softdep directive ensures vfio-pci loads before the ism module without completely blocking ism from the host. The ism module remains available in the kernel but does not own the ISM device.
Prerequisites
- You have installed OpenShift Container Platform 4.21 or later.
- You have installed the OpenShift Virtualization Operator.
- You have cluster administrator permissions.
- You have installed the Butane tool for generating Ignition-compatible
MachineConfigmanifests. - The ISM virtual PCI device is present and visible on the PCI bus of the target nodes.
- You have installed the OpenShift CLI (
oc).
Procedure
-
On each node, confirm the ISM device is displayed as a PCI device by running the following command:
-
Retrieve the PCI vendor and device ID by running the following command:
Record the combined PCI vendor and device ID
1014:04ed. This value is used in thevfio-pciandHyperConvergedconfiguration. -
Create a Butane configuration file named
100-master-vfio-ism.buto bind the ISM device tovfio-pci:variant: openshift version: 4.22.0 metadata: name: 100-master-vfio-ism labels: machineconfiguration.openshift.io/role: master storage: files: - path: /etc/modprobe.d/vfio-ism.conf mode: 0644 overwrite: true contents: inline: | softdep ism pre: vfio-pci options vfio-pci ids=1014:04ed openshift: kernel_arguments: - rd.driver.pre=vfio-pci - vfio-pci.ids=1014:04edwhere:
storage.files[].contents.inline softdep ism pre: vfio-pci- Specifies that
vfio-pcimust load before theismmodule. Theismmodule remains available on the host but does not own the device. storage.files[].contents.inline options vfio-pci ids=1014:04ed- Specifies that
vfio-pciclaims devices with this PCI vendor and device ID. openshift.kernel_arguments rd.driver.pre=vfio-pci- Specifies that
vfio-pciloads during initramfs before any other driver. openshift.kernel_arguments vfio-pci.ids=1014:04ed- Specifies the PCI device ID passed directly to
vfio-pciat boot time.
-
Convert the Butane file to a
MachineConfigmanifest by running the following command: -
Apply the
MachineConfigto the cluster by running the following command: -
Watch the
MachineConfigrollout and wait for completion before proceeding: -
Confirm the ISM device resource is visible and allocatable on the nodes by running the following command:
-
Edit the
HyperConvergedcustom resource to expose the ISM device by running the following command:Add the ISM device under
spec.virtualization.permittedHostDevices: -
Verify the HyperConverged Operator accepted the configuration by running the following command:
-
Add the ISM device to a
VirtualMachinemanifest: -
Apply the
VirtualMachinemanifest by running the following command:
Verification
-
Verify
vfio-pcibinding on all control plane nodes by running the following command:$ for node in $(oc get nodes -l node-role.kubernetes.io/master -o name); do echo "=== $node ===" oc debug $node -- chroot /host bash -c \ "lspci -nnk | grep -A2 'ISM\|1014:04ed'" 2>/dev/null doneExample output=== node/master-0 === 0000:00:00.0 Non-VGA unclassified device [0000]: IBM Internal Shared Memory (ISM) virtual PCI device [1014:04ed] Kernel driver in use: vfio-pci Kernel modules: ism === node/master-1 === 0000:00:00.0 Non-VGA unclassified device [0000]: IBM Internal Shared Memory (ISM) virtual PCI device [1014:04ed] Kernel driver in use: vfio-pci Kernel modules: ism === node/master-2 === 0000:00:00.0 Non-VGA unclassified device [0000]: IBM Internal Shared Memory (ISM) virtual PCI device [1014:04ed] Kernel driver in use: vfio-pci Kernel modules: ismKernel modules: ismindicates theismmodule is available in the kernel but is not actively managing the device.vfio-pciowns the device. -
Verify the ISM device is present inside the VM by connecting to the VM console and running the following command:
Example output0001:00:00.0 Non-VGA unclassified device: IBM Internal Shared Memory (ISM) virtual PCI deviceThe output confirms PCI passthrough. The guest binds the
ismdriver only if the guest image includes that module.
Additional resources