Configuring PCI passthrough
The Peripheral Component Interconnect (PCI) passthrough feature enables you to access and manage hardware devices from a virtual machine (VM). When PCI passthrough is configured, the PCI devices function as if they were physically attached to the guest operating system.
Cluster administrators can expose and manage host devices that are permitted to be used in the cluster by using the oc command-line interface (CLI).
For vfio-pci to allocate a PCI device, no other kernel driver can manage that device. If a driver already manages the device, you must add the specific kernel module to a blocklist.
Adding a kernel module to a blocklist makes all devices handled by that module unavailable to the host.
The following example shows a MachineConfig CR that adds the enic network driver to a blocklist by creating a configuration file in /etc/modprobe.d/ and adding kernel arguments:
apiVersion: machineconfiguration.openshift.io/v1
kind: MachineConfig
metadata:
labels:
machineconfiguration.openshift.io/role: worker
name: 100-blacklist-enic
spec:
config:
ignition:
version: 3.4.0
storage:
files:
- contents:
source: data:,blacklist%20enic%0A
mode: 420
overwrite: true
path: /etc/modprobe.d/blacklist-enic.conf
kernelArguments:
- enic.blacklist=1
- rd.driver.blacklist=enicPreparing nodes for GPU passthrough
You can prevent GPU operands from deploying on worker nodes that you designated for GPU passthrough.
Preventing NVIDIA GPU operands from deploying on nodes
If you use the NVIDIA GPU Operator in your cluster, you can apply the nvidia.com/gpu.deploy.operands=false label to nodes that you do not want to configure for GPU or vGPU operands. This prevents the creation of the pods that configure GPU or vGPU operands and terminates existing pods.
Prerequisites
- The OpenShift CLI (
oc) is installed.
Procedure
- Label the node by running the following command:terminal
$ oc label node <node_name> nvidia.com/gpu.deploy.operands=falsewhere:
<node_name>Specifies the name of a node where you do not want to install the NVIDIA GPU operands.
Verification
- Verify that the label was added to the node by running the following command:terminal
$ oc describe node <node_name> - Optional: If GPU operands were previously deployed on the node, verify their removal.
- Check the status of the pods in the
nvidia-gpu-operatornamespace by running the following command:terminal$ oc get pods -n nvidia-gpu-operatorExample output:
terminalNAME READY STATUS RESTARTS AGE gpu-operator-59469b8c5c-hw9wj 1/1 Running 0 8d nvidia-sandbox-validator-7hx98 1/1 Running 0 8d nvidia-sandbox-validator-hdb7p 1/1 Running 0 8d nvidia-sandbox-validator-kxwj7 1/1 Terminating 0 9d nvidia-vfio-manager-7w9fs 1/1 Running 0 8d nvidia-vfio-manager-866pz 1/1 Running 0 8d nvidia-vfio-manager-zqtck 1/1 Terminating 0 9d - Monitor the pod status until the pods with
Terminatingstatus are removed:terminal$ oc get pods -n nvidia-gpu-operatorExample output:
terminalNAME READY STATUS RESTARTS AGE gpu-operator-59469b8c5c-hw9wj 1/1 Running 0 8d nvidia-sandbox-validator-7hx98 1/1 Running 0 8d nvidia-sandbox-validator-hdb7p 1/1 Running 0 8d nvidia-vfio-manager-7w9fs 1/1 Running 0 8d nvidia-vfio-manager-866pz 1/1 Running 0 8d
- Check the status of the pods in the
Preparing host devices for PCI passthrough
About preparing a host device for PCI passthrough
To prepare a host device for PCI passthrough by using the CLI, create a MachineConfig object and add kernel arguments to enable the Input-Output Memory Management Unit (IOMMU).
Bind the PCI device to the Virtual Function I/O (VFIO) driver and then expose it in the cluster by editing the permittedHostDevices field of the HyperConverged custom resource (CR). The permittedHostDevices list is empty when you first install the OpenShift Virtualization Operator.
To remove a PCI host device from the cluster by using the CLI, delete the PCI device information from the HyperConverged CR.
Adding kernel arguments to enable the IOMMU driver
You must enable the Input-Output Memory Management Unit (IOMMU) driver before you can configure mediated devices. To enable the IOMMU driver in the kernel, create the MachineConfig object and add the kernel arguments.
Prerequisites
- You have cluster administrator permissions.
- Your CPU hardware is Intel or AMD.Note
Enabling IOMMU is not required on
s390xarchitecture. - You enabled Intel Virtualization Technology for Directed I/O extensions or AMD IOMMU in the BIOS.
- You have installed the OpenShift CLI (
oc).
Procedure
- Create a
MachineConfigobject that identifies the kernel argument. The following example shows a kernel argument for an Intel CPU.yamlapiVersion: machineconfiguration.openshift.io/v1 kind: MachineConfig metadata: labels: machineconfiguration.openshift.io/role: worker name: 100-worker-iommu spec: config: ignition: version: 3.2.0 kernelArguments: - intel_iommu=on # ...metadata.labels.machineconfiguration.openshift.io/rolespecifies that the new kernel argument is applied only to worker nodes.metadata.namespecifies the ranking of this kernel argument (100) among the machine configs and its purpose. If you have an AMD CPU, specify the kernel argument asamd_iommu=on.spec.kernelArgumentsspecifies the kernel argument asintel_iommufor an Intel CPU.
- Create the new
MachineConfigobject:terminal$ oc create -f 100-worker-kernel-arg-iommu.yaml
Verification
- Verify that the new
MachineConfigobject was added by entering the following command and observing the output:terminal$ oc get MachineConfigExample output:
terminalNAME IGNITIONVERSION AGE 00-master 3.5.0 164m 00-worker 3.5.0 164m 01-master-container-runtime 3.5.0 164m 01-master-kubelet 3.5.0 164m 01-worker-container-runtime 3.5.0 164m 01-worker-kubelet 3.5.0 164m 100-master-chrony-configuration 3.5.0 169m 100-master-set-core-user-password 3.5.0 169m 100-worker-chrony-configuration 3.5.0 169m 100-worker-iommu 3.5.0 14s - Verify that IOMMU is enabled at the operating system (OS) level by entering the following command:terminal
$ dmesg | grep -i iommu- If IOMMU is enabled, output is displayed as shown in the following example:
Example output:
terminalIntel: [ 0.000000] DMAR: Intel(R) IOMMU Driver AMD: [ 0.000000] AMD-Vi: IOMMU Initialized
- If IOMMU is enabled, output is displayed as shown in the following example:
Binding PCI devices to the VFIO driver
To bind PCI devices to the VFIO (Virtual Function I/O) driver, obtain the values for the vendor ID and the device ID from each device and create a list with the values. Add this list to the MachineConfig object.
The MachineConfig Operator generates the /etc/modprobe.d/vfio.conf on the nodes with the PCI devices, and binds the PCI devices to the VFIO driver.
Prerequisites
- You added kernel arguments to enable IOMMU for the CPU.Note
Enabling IOMMU is not required on
s390xarchitecture. - You have installed the OpenShift CLI (
oc).
Procedure
- Run the
lspcicommand with the name of the GPU accelerator to obtain the vendor ID and the device ID for the PCI device.NoteNVIDIA GPU is supported on
x86andaarch64architectures, Intel QAT is supported onx86architecture, and IBM(R) Spyre is supported ons390xarchitecture.terminal$ lspci -nnv | grep -i <gpu_accelerator>Valid values for
<gpu_accelerator>arenvidia,qat, andspyre.Example output:
terminal02:01.0 3D controller [0302]: NVIDIA Corporation GV100GL [Tesla V100 PCIe 32GB] [10de:1eb8] (rev a1) - Create a Butane config file,
100-worker-vfiopci.bu, binding the PCI device to the VFIO driver.NoteThe Butane version you specify in the config file should match the OpenShift Container Platform version and always ends in
0. For example,4.22.0. See "Creating machine configs with Butane" for information about Butane.Example:
yamlvariant: openshift version: 4.22.0 metadata: name: 100-worker-vfiopci labels: machineconfiguration.openshift.io/role: worker storage: files: - path: /etc/modprobe.d/vfio.conf mode: 0644 overwrite: true contents: inline: | options vfio-pci ids=<vendor_id>:<device_id> - path: /etc/modules-load.d/vfio-pci.conf mode: 0644 overwrite: true contents: inline: vfio-pcimetadata.labels.machineconfiguration.openshift.io/role: workerspecifies that the new kernel argument is applied only to compute nodes.storage.files.contents.inline, where the path is/etc/modprobe.d/vfio.conf, specifies the previously determined hexadecimal vendor ID and device ID values to bind a device to the VFIO driver. You can add a list of multiple devices with their vendor and device information.storage.files.path, where thecontents.inlineisvfio-pci, specifies the file that loads thevfio-pcikernel module on the compute nodes.
- Use Butane to generate a
MachineConfigobject file,100-worker-vfiopci.yaml, containing the configuration to be delivered to the compute nodes:terminal$ butane 100-worker-vfiopci.bu -o 100-worker-vfiopci.yaml - Apply the
MachineConfigobject to the compute nodes:terminal$ oc apply -f 100-worker-vfiopci.yaml - Verify that the
MachineConfigobject was added.terminal$ oc get MachineConfigExample output:
terminalNAME GENERATEDBYCONTROLLER IGNITIONVERSION AGE 00-master d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h 00-worker d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h 01-master-container-runtime d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h 01-master-kubelet d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h 01-worker-container-runtime d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h 01-worker-kubelet d3da910bfa9f4b599af4ed7f5ac270d55950a3a1 3.5.0 25h 100-worker-iommu 3.5.0 30s 100-worker-vfiopci-configuration 3.5.0 30s
Verification
- Verify that the VFIO driver is loaded.terminal
$ lspci -nnk -d <vendor_id>:The output confirms that the VFIO driver is being used.
Example output:
04:00.0 3D controller [0302]: NVIDIA Corporation GP102GL [Tesla P40] [10de:1eb8] (rev a1) Subsystem: NVIDIA Corporation Device [10de:1eb8] Kernel driver in use: vfio-pci Kernel modules: nouveau
Exposing PCI host devices in the cluster using the CLI
To expose PCI host devices in the cluster, add details about the PCI devices to the spec.permittedHostDevices.pciHostDevices array of the HyperConverged custom resource (CR).
Prerequisites
- You have installed the OpenShift CLI (
oc).
Procedure
- Edit the
HyperConvergedCR in your default editor by running the following command:terminal$ oc edit hyperconvergeds.v1beta1.hco.kubevirt.io kubevirt-hyperconverged -n openshift-cnv - Add the PCI device information to the
spec.permittedHostDevices.pciHostDevicesarray.Example configuration file:
yamlapiVersion: hco.kubevirt.io/v1beta1 kind: HyperConverged metadata: name: kubevirt-hyperconverged namespace: openshift-cnv spec: permittedHostDevices: pciHostDevices: - pciDeviceSelector: "10DE:1DB6" resourceName: "nvidia.com/GV100GL_Tesla_V100" - pciDeviceSelector: "10DE:1EB8" resourceName: "nvidia.com/TU104GL_Tesla_T4" - pciDeviceSelector: "8086:6F54" resourceName: "intel.com/qat" externalResourceProvider: true # ...spec.permittedHostDevicesspecifies the host devices that are permitted to be used in the cluster.spec.permittedHostDevices.pciHostDevicesspecifies the list of PCI devices available on the node.spec.permittedHostDevices.pciHostDevices.pciDeviceSelectorspecifies the vendor ID and the device ID required to identify the PCI device.spec.permittedHostDevices.pciHostDevices.resourceNamespecifies the name of a PCI host device.spec.permittedHostDevices.pciHostDevices.externalResourceProvideris an optional setting. Setting this field totrueindicates that the resource is provided by an external device plugin. OpenShift Virtualization allows the usage of this device in the cluster but leaves the allocation and monitoring to an external device plugin.NoteThe above example snippet shows two PCI host devices that are named
nvidia.com/GV100GL_Tesla_V100andnvidia.com/TU104GL_Tesla_T4added to the list of permitted host devices in theHyperConvergedCR. These devices have been tested and verified to work with OpenShift Virtualization.Example configuration file for an IBM(R) Spyre device on
s390xarchitecture:yamlapiVersion: hco.kubevirt.io/v1beta1 kind: HyperConverged metadata: name: kubevirt-hyperconverged namespace: openshift-cnv spec: permittedHostDevices: pciHostDevices: - pciDeviceSelector: "1014:06a8" resourceName: "ibm.com/spyre" # ...
- Save your changes and exit the editor.
Verification
- Verify that the PCI host devices were added to the node by running the following command. The example output shows that there is one device each associated with the
nvidia.com/GV100GL_Tesla_V100,nvidia.com/TU104GL_Tesla_T4, andintel.com/qatresource names.terminal$ oc describe node <node_name>Example output:
terminalCapacity: cpu: 64 devices.kubevirt.io/kvm: 110 devices.kubevirt.io/tun: 110 devices.kubevirt.io/vhost-net: 110 ephemeral-storage: 915128Mi hugepages-1Gi: 0 hugepages-2Mi: 0 memory: 131395264Ki nvidia.com/GV100GL_Tesla_V100 1 nvidia.com/TU104GL_Tesla_T4 1 intel.com/qat: 1 pods: 250 Allocatable: cpu: 63500m devices.kubevirt.io/kvm: 110 devices.kubevirt.io/tun: 110 devices.kubevirt.io/vhost-net: 110 ephemeral-storage: 863623130526 hugepages-1Gi: 0 hugepages-2Mi: 0 memory: 130244288Ki nvidia.com/GV100GL_Tesla_V100 1 nvidia.com/TU104GL_Tesla_T4 1 intel.com/qat: 1 pods: 250NoteWhen using an IBM(R) Spyre device on
s390xarchitecture, the allocated device is shown as follows:ibm.com/spyre: 1.
Removing PCI host devices from the cluster using the CLI
To remove a PCI host device from the cluster, delete the information for that device from the HyperConverged custom resource (CR).
Prerequisites
- You have installed the OpenShift CLI (
oc).
Procedure
- Edit the
HyperConvergedCR in your default editor by running the following command:terminal$ oc edit hyperconvergeds.v1beta1.hco.kubevirt.io kubevirt-hyperconverged -n openshift-cnv - Remove the PCI device information from the
spec.permittedHostDevices.pciHostDevicesarray by deleting thepciDeviceSelector,resourceNameandexternalResourceProvider(if applicable), fields for the appropriate device. In this example, the user deletes thenvidia.com/TU104GL_Tesla_T4.Example configuration file:
yamlapiVersion: hco.kubevirt.io/v1beta1 kind: HyperConverged metadata: name: kubevirt-hyperconverged namespace: openshift-cnv spec: permittedHostDevices: pciHostDevices: - pciDeviceSelector: "10DE:1DB6" resourceName: "nvidia.com/GV100GL_Tesla_V100" # ... - Save your changes and exit the editor.
Verification
- Verify that you removed the PCI host device from the node by running the following command. The example output shows that there are zero devices associated with the
nvidia.com/TU104GL_Tesla_T4resource name.terminal$ oc describe node <node_name>Example output:
terminalCapacity: cpu: 64 devices.kubevirt.io/kvm: 110 devices.kubevirt.io/tun: 110 devices.kubevirt.io/vhost-net: 110 ephemeral-storage: 915128Mi hugepages-1Gi: 0 hugepages-2Mi: 0 memory: 131395264Ki nvidia.com/GV100GL_Tesla_V100 1 nvidia.com/TU104GL_Tesla_T4 0 pods: 250 Allocatable: cpu: 63500m devices.kubevirt.io/kvm: 110 devices.kubevirt.io/tun: 110 devices.kubevirt.io/vhost-net: 110 ephemeral-storage: 863623130526 hugepages-1Gi: 0 hugepages-2Mi: 0 memory: 130244288Ki nvidia.com/GV100GL_Tesla_V100 1 nvidia.com/TU104GL_Tesla_T4 0 pods: 250
Configuring virtual machines for PCI passthrough
After the PCI devices have been added to the cluster, you can assign them to virtual machines. The PCI devices are now available as if they are physically connected to the virtual machines.
Assigning a PCI device to a virtual machine
When a PCI device is available in a cluster, you can assign it to a virtual machine and enable PCI passthrough.
Procedure
- Assign the PCI device to a virtual machine as a host device.
Example:
yamlapiVersion: kubevirt.io/v1 kind: VirtualMachine spec: domain: devices: hostDevices: - deviceName: nvidia.com/TU104GL_Tesla_T4 name: hostdevices1spec.template.spec.domain.devices.hostDevices.deviceNamespecifies the name of the PCI device that is permitted on the cluster as a host device. The virtual machine can access this host device. When using an IBM(R) Spyre device ons390xarchitecture, specifyibm.com/spyre:.
Verification
- Use the following command to verify that the host device is available from the virtual machine.terminal
$ lspci -nnk | grep <gpu_accelerator>Valid values for
<gpu_accelerator>arenvidia,qat, andspyre.Example output:
terminal$ 02:01.0 3D controller [0302]: NVIDIA Corporation GV100GL [Tesla V100 PCIe 32GB] [10de:1eb8] (rev a1)