Add worker nodes and provision DPUs
After the DPF Operator and the hosted cluster are configured, adjust the OVN-Kubernetes CNI settings, add DPU-equipped worker nodes to the management cluster, and provision the DPUs.
Enable the OVN-Kubernetes resource injector
You can install the OVN-Kubernetes resource injector by using Helm to deploy a mutating admission webhook that automatically injects SR-IOV virtual function resource requests and network attachment annotations into each pod scheduled to a worker node.
Virtual function resource capacity on worker nodes is provided by the NodeSRIOVDevicePluginConfig resource, which replaces the manual SR-IOV device plugin DaemonSet and control plane node patching used in earlier DPF versions.
Prerequisites
- You have access to the management cluster as a user with the
cluster-adminrole. - You have installed the
ocCLI. - You have installed the
helmCLI. - You have set the DPF Operator environment variables. For details, see "DPF Operator installation environment variables".
- You have created the
NodeSRIOVDevicePluginConfigresource.
Procedure
-
Install the OVN-Kubernetes resource injector by using Helm:
$ helm upgrade --install -n openshift-ovn-kubernetes ovn-kubernetes \"$OVN_TEMPLATE_CHART_URL/ovn-kubernetes-chart" \--version "${OVN_CHART_VERSION}" \--skip-crds \--set ovn-kubernetes-resource-injector.enabled=true \--set ovn-kubernetes-resource-injector.resourceName="openshift.io/bf3_vfs" \--set ovn-kubernetes-resource-injector.prioritizeOffloading=false \--set ovn-kubernetes-resource-injector.controllerManager.hostNetwork=true \--set ovn-kubernetes-resource-injector.controllerManager.webhookPort="19443" \--set ovn-kubernetes-resource-injector.controllerManager.healthProbeBindAddress=":18081" \--set ovn-kubernetes-resource-injector.controllerManager.webhook.image.pullPolicy=IfNotPresent \--set "ovn-kubernetes-resource-injector.controllerManager.webhook.args={--leader-elect,--metrics-bind-address=:29091}" \--set nodeWithDPUManifests.enabled=false \--set nodeWithoutDPUManifests.enabled=false \--set dpuManifests.enabled=false \--set controlPlaneManifests.enabled=false \--set commonManifests.enabled=falseExample outputNAME: ovn-kubernetesLAST DEPLOYED: Sun Nov 2 17:10:29 2025NAMESPACE: openshift-ovn-kubernetesSTATUS: deployedREVISION: 1DESCRIPTION: Install completeTEST SUITE: None
Verification
-
Verify the resource injector mutating webhook configuration was applied:
$ oc get mutatingwebhookconfiguration | grep ovnExample outputNAME WEBHOOKS AGEovn-kubernetes-ovn-kubernetes-resource-injector 1 22h
OVN-Kubernetes DPU-Host mode
DPU-Host mode on worker nodes with accelerated OVN-Kubernetes CNI is automatically configured by the DPF provisioning controller.
When the DPFOperatorConfig resource is created and worker nodes with the worker-dpu label are provisioned, the DPF provisioning controller automatically configures the required settings for DPU-Host mode, including the network node identity and hardware offload configuration.
Add worker nodes by using the Bare Metal Operator
You can add DPU-equipped worker nodes to the management cluster by creating BareMetalHost resources that the Bare Metal Operator provisions.
Prerequisites
- You have access to the management cluster as a user with the
cluster-adminrole. - You have installed the
ocCLI. - You have installed the Bare Metal Operator on the management cluster.
- Physical worker servers with Redfish-compatible BMC, iDRAC, or iLO access are available.
- Network connectivity exists from the management cluster to the worker BMC interfaces.
- You have the BMC IP address and access credentials for each server.
- You have the MAC address of the management network interface for each server.
- You have the name of the root disk device for each server.
Procedure
-
Set the following environment variables for the worker node:
$ export BMC_IP=<bmc_ip_address>$ export BMC_USER=<bmc_username>$ export BMC_PASSWORD=<bmc_password>$ export WORKER_NAME=<worker_name>$ export BOOT_MAC=<management_interface_mac>$ export ROOT_DEVICE=<root_device_path>where:
<bmc_ip_address>- Specifies the IP address of the worker node BMC interface.
<bmc_username>- Specifies the username for BMC access.
<bmc_password>- Specifies the password for BMC access.
<worker_name>- Specifies a name for the worker node, such as
worker-01. <management_interface_mac>- Specifies the MAC address of the out-of-band management interface, such as
00:00:5E:00:53:01. <root_device_path>- Specifies the path to the root disk device, such as
/dev/nvme0n1.
-
Verify BMC connectivity from one of the control plane nodes:
$ ping $BMC_IP$ curl -k https://$BMC_IP/redfish/v1/$ curl -k -u $BMC_USER:$BMC_PASSWORD https://$BMC_IP/redfish/v1/Systems -
Verify that the Bare Metal Operator is available:
$ oc get clusteroperator baremetal -
Create a file named
provisioning.yamlwith the following content to disable the provisioning network:apiVersion: metal3.io/v1alpha1kind: Provisioningmetadata:name: provisioning-configurationspec:provisioningNetwork: "Disabled"watchAllNamespaces: falsewarningWhen
provisioningNetworkis set toDisabled, servers boot by using Redfish virtual media instead of PXE. -
Apply the
Provisioningresource:$ oc apply -f provisioning.yaml -
Create a file named
bmc-secret.yamlwith the following content to store the BMC credentials:apiVersion: v1kind: Secretmetadata:name: ${WORKER_NAME}-bmc-secretnamespace: openshift-machine-apitype: OpaquestringData:username: ${BMC_USER}password: ${BMC_PASSWORD} -
Apply the BMC credentials secret:
$ envsubst < bmc-secret.yaml | oc apply -f - -
Create a file named
baremetalhost.yaml. TheuserDatasecret determines the node type:-
For a DPU-equipped worker node, reference the
worker-dpu-user-data-managedsecret:apiVersion: metal3.io/v1alpha1kind: BareMetalHostmetadata:name: ${WORKER_NAME}namespace: openshift-machine-apispec:online: truebootMACAddress: ${BOOT_MAC}rootDeviceHints:deviceName: ${ROOT_DEVICE}bmc:address: redfish-virtualmedia+https://${BMC_IP}credentialsName: ${WORKER_NAME}-bmc-secretdisableCertificateVerification: truecustomDeploy:method: install_coreosuserData:name: worker-dpu-user-data-managednamespace: openshift-machine-api -
For a regular worker node without a DPU, reference the
worker-user-data-managedsecret instead:apiVersion: metal3.io/v1alpha1kind: BareMetalHostmetadata:name: ${WORKER_NAME}namespace: openshift-machine-apispec:online: truebootMACAddress: ${BOOT_MAC}rootDeviceHints:deviceName: ${ROOT_DEVICE}bmc:address: redfish-virtualmedia+https://${BMC_IP}credentialsName: ${WORKER_NAME}-bmc-secretdisableCertificateVerification: truecustomDeploy:method: install_coreosuserData:name: worker-user-data-managednamespace: openshift-machine-apiwarningAdding a regular worker node without a DPU is a Technology Preview feature.
-
-
Apply the
BareMetalHostresource:$ envsubst < baremetalhost.yaml | oc apply -f -
Verification
-
Monitor the provisioning progress:
$ oc get bmh -n openshift-machine-api -wExample outputNAME STATE CONSUMER ONLINE ERROR AGEworker-01 registering true 10sworker-01 inspecting true 15sworker-01 preparing true 20sworker-01 available true 30sworker-01 provisioning true 1mworker-01 provisioned true 10m
Approve worker node CSRs
You must approve the pending certificate signing requests (CSRs) for worker nodes that join the management cluster.
Worker nodes provisioned by using a BareMetalHost resource do not have an associated Machine object, so the default OpenShift machine approver does not automatically approve their certificate signing requests (CSRs). You must manually approve the kube-apiserver-client-kubelet CSR from the node-bootstrapper service account and the kubelet-serving CSR from the node for each worker node.
Prerequisites
- You have access to the management cluster as a user with the
cluster-adminrole. - You have installed the
ocCLI. - Worker nodes are booted and attempting to join the management cluster.
Procedure
-
Watch for pending CSRs:
$ oc get csr -w -
Approve all pending CSRs:
$ oc get csr -o go-template='{{range .items}}{{if not .status}}{{.metadata.name}}{{"\n"}}{{end}}{{end}}' | xargs oc adm certificate approveExample outputcertificatesigningrequest.certificates.k8s.io/csr-27bgq approvedcertificatesigningrequest.certificates.k8s.io/csr-69g65 approvedcertificatesigningrequest.certificates.k8s.io/csr-7r862 approvedcertificatesigningrequest.certificates.k8s.io/csr-f5vk7 approvedRepeat this step until no pending CSRs remain. Each node typically generates multiple CSRs.
-
Verify that the worker nodes joined the cluster:
$ oc get nodesExample outputNAME STATUS ROLES AGE VERSIONhost-worker1 NotReady worker 68s v1.35.6host-worker2 NotReady worker 75s v1.35.6master-0 Ready control-plane,master,worker 4d22h v1.35.6master-1 Ready control-plane,master,worker 4d21h v1.35.6master-2 Ready control-plane,master,worker 4d22h v1.35.6noteThe worker nodes show a status of
NotReadyuntil the DPU provisioning process is fully completed and all OVN-Kubernetes CNI components on the host and the DPU are running. Do not proceed to the next steps until all pending CSRs are approved.
Verify DPU provisioning
After the worker nodes join the management cluster, the DPUSet controller automatically detects nodes with the feature.node.kubernetes.io/dpu-enabled label, which the Node Feature Discovery Operator applies to DPU-equipped nodes. The controller then creates a DPU object for each node and starts the provisioning process. You can monitor the provisioning stages to verify progress.
Prerequisites
- You have access to the management cluster as a user with the
cluster-adminrole. - You have installed the
ocCLI. - The worker node CSRs are approved and the nodes have joined the management cluster.
Procedure
-
Watch for
DPUobject creation:$ oc get dpu -n dpf-operator-system -wExample outputNAME READY OPERATIONAL PHASE AGE<node-name>-<dpu-id> Unknown Node Effect 25s<node-name>-<dpu-id> Unknown Initialize Interface 26s<node-name>-<dpu-id> Unknown Config FW Parameters 28s<node-name>-<dpu-id> Unknown Prepare BFB 28s<node-name>-<dpu-id> Unknown OS Installing 5m28s<node-name>-<dpu-id> Unknown DPU Config 15m<node-name>-<dpu-id> Unknown Rebooting 26m<node-name>-<dpu-id> Unknown Host Network Configuration 27m<node-name>-<dpu-id> False DPU Cluster Config 64m<node-name>-<dpu-id> Unknown Node Effect Removal 71m<node-name>-<dpu-id> True True Ready 71mThe
DPUobjects progress through the following provisioning stages:Initializing- The
DPUobject is created. OS Installing- The BFB installation is in progress.
Rebooting- The host and DPU are resetting.
DPU Cluster Config- The DPU Kubernetes node join procedure is in progress. DPU CSRs are automatically approved by the DPF HCP Provisioner Operator.
Host Network Configuration- Networking configuration adjustments are applied on the host.
Ready- The DPU is successfully provisioned and ready to use.
Error- Provisioning failed. Check events and conditions for details.
warningWhen the provisioning stage reaches
DPU Cluster Config, proceed to "Configure authorization for the hosted cluster" to complete the DPU node join process. -
Monitor detailed provisioning progress:
$ oc -n dpf-operator-system exec deploy/dpf-operator-controller-manager -- /dpfctl describe dpudeployments -
Optional: View detailed status for a specific
DPUobject: In the following command, replace<dpu_name>with the name of theDPUresource:$ oc describe dpu -n dpf-operator-system <dpu_name> -
Optional: Follow the provisioning controller logs for a specific DPU: In the following command, replace
<dpu_name>with the name of theDPUresource:$ oc logs -n dpf-operator-system -l dpu.nvidia.com/component=dpf-provisioning-controller-manager --tail=-1 -f | grep <dpu_name>
Configure authorization for the hosted cluster
DPF services running on DPU nodes require privileged access to host networking and devices. You must create a ClusterRoleBinding on the hosted cluster that grants the privileged security context constraint (SCC) to all service accounts in the dpf-operator-system namespace.
Prerequisites
- You have access to the hosted cluster as a user with the
cluster-adminrole. - You have installed the
ocCLI. - The hosted cluster kubeconfig file is available.
- DPU provisioning has reached the
DPU Cluster Configstage.
Procedure
- Get the hosted cluster kubeconfig:
$ oc get secret $HOSTED_CLUSTER_NAME-admin-kubeconfig -n $CLUSTERS_NAMESPACE -o jsonpath='{.data.kubeconfig}' | base64 -d > $HOSTED_CLUSTER_NAME.kubeconfig
- Switch to the hosted cluster context:
$ export KUBECONFIG="$(pwd)/$HOSTED_CLUSTER_NAME.kubeconfig"
- Create a file named
dpu-cluster-scc.yamlwith the following content:apiVersion: rbac.authorization.k8s.io/v1kind: ClusterRoleBindingmetadata:name: dpf-system-scc-privilegedlabels:app.kubernetes.io/component: rbacapp.kubernetes.io/part-of: dpu-servicesroleRef:apiGroup: rbac.authorization.k8s.iokind: ClusterRolename: system:openshift:scc:privilegedsubjects:- kind: GroupapiGroup: rbac.authorization.k8s.ioname: system:serviceaccounts:dpf-operator-system - Apply the resource file on the hosted cluster:
$ oc apply -f dpu-cluster-scc.yaml
Verify full system readiness
After the DPU provisioning process completes, you can verify that all worker nodes, SR-IOV virtual functions, and DPU services are operational on the management cluster.
Prerequisites
- You have access to the management cluster as a user with the
cluster-adminrole. - You have installed the
ocCLI. - DPU provisioning has completed.
Procedure
-
Switch back to the management cluster context:
$ export KUBECONFIG="$(pwd)/mgmt-kubeconfig" -
Verify that all worker nodes are in a
Readystate:$ oc get nodesExample outputNAME STATUS ROLES AGE VERSIONhost-worker1 Ready worker 57m v1.35.6host-worker2 Ready worker 57m v1.35.6master-0 Ready control-plane,master,worker 4d23h v1.35.6master-1 Ready control-plane,master,worker 4d22h v1.35.6master-2 Ready control-plane,master,worker 4d23h v1.35.6 -
Verify that SR-IOV virtual functions are registered as Kubernetes node resources on the worker nodes:
$ oc get nodes -l 'node-role.kubernetes.io/worker,!node-role.kubernetes.io/control-plane' -o json | \jq '.items[] | {name: .metadata.name, capacity: .status.capacity."openshift.io/bf3_vfs", allocatable: .status.allocatable."openshift.io/bf3_vfs"}'Example output{"name": "host-worker1","capacity": "90","allocatable": "90"}{"name": "host-worker2","capacity": "90","allocatable": "90"} -
Verify that all DPU services are in a
Successphase:$ oc get dpuservices -n dpf-operator-systemExample outputNAME READY PHASE AGEdoca-telemetry-service-7s8pb True Success 42mflannel True Success 26hhbn-gffmv True Success 25mkube-state-metrics-rbac True Success 4h10mnode-problem-detector True Success 4h10mnvidia-k8s-ipam-node True Success 4h10movn-f49zx True Success 17movs-cni True Success 26hservicechainset-rbac-and-crds True Success 138msfc-controller True Success 26hsriov-device-plugin True Success 26h -
Optional: View detailed DPU service status:
$ oc -n dpf-operator-system exec deploy/dpf-operator-controller-manager -- /dpfctl describe all --show-resources=dpuservice --grouping=false