Troubleshoot DPF
You can diagnose and resolve common NVIDIA DPF Operator issues with DPU provisioning, hosted cluster readiness, networking, and collect diagnostic logs for support. These procedures complement the official NVIDIA debugging tools and guides.
DPU provisioning does not start
If DPU provisioning does not start immediately after you add worker nodes to the management cluster, verify that certificate signing requests (CSRs), controller pods, Node Feature Discovery (NFD) labels, and DPF resource objects are in the correct state.
- Verify that all worker CSRs are approved
- Run the following command to list the CSR status on the management cluster:
Ensure that all CSRs for the worker nodes show an$ oc get csr
Approvedstatus. - Verify that all DPF controller pods are running
- Run the following command to check the status of the DPF Operator pods:
Ensure that all pods are in a$ oc get pod -n dpf-operator-system
Runningstate. - Verify that worker nodes are labeled for DPU provisioning by NFD
- Run the following command to confirm that the
dpu-enabledlabel is present on the worker nodes:The output lists the worker nodes that NFD has labeled for DPU provisioning. For example:$ oc get nodes -l feature.node.kubernetes.io/dpu-enabled=""Example outputNAME STATUS ROLES AGE VERSIONhost-worker1 NotReady worker,worker-dpu 62s v1.35.6host-worker2 NotReady worker,worker-dpu 66s v1.35.6 - Check BFB object status
- Run the following command to verify that the BlueField Bootstream File (BFB) image is downloaded and ready:
Check the$ oc describe bfb -n dpf-operator-system bf-bundle
status.conditionsfield for download progress and any error messages. - Verify that the BFB image URL is reachable
- If the BFB download fails, run the following command to confirm that the image URL is reachable, replacing
$BFB_URLwith the image URL:Ensure that the response returns a$ curl -I $BFB_URL200 OKstatus code. If the download fails, verify network connectivity to the image registry, check for firewall or proxy restrictions, and ensure that sufficient disk space is available on the node. - Monitor DPU provisioning progress
- Run the following command to watch the DPU objects progress through provisioning:
Wait for each DPU to progress from$ oc get dpu -n dpf-operator-system -w
PendingtoProvisioningtoReady. - Verify that the DPU hardware is detected on the worker node
- Open a debug shell on the worker node:
Inside the debug shell, run the following commands to confirm that a BlueField device is present:$ oc debug node/<worker-node-name>sh-5.1# chroot /hostIf the DPU is not listed, verify that it is properly seated in the PCIe slot and that it is not disabled in the server BIOS.sh-5.1# lspci | grep -i mellanox
- Verify DPU firmware and software compatibility
- Inside the debug shell, run the following command to check the DPU firmware version:
Confirm that the BlueField firmware and DOCA software versions on the DPU are compatible with the DPF Operator version that you deployed.sh-5.1# mlxfwmanager --query
- Check
DPUDeploymentobject status - Inspect the
DPUDeploymentobject for information about the following resources:- BFB object state
DPUServiceTemplateobjects stateDPUServiceConfigurationobjects state
DPUDeploymentstatus:Alternatively, run the following$ oc get dpudeployments -n dpf-operator-system dpudeployment -o yamldpfctlcommand for a summarized view:$ oc -n dpf-operator-system exec deploy/dpf-operator-controller-manager -- /dpfctl describe dpudeployments
DPU objects remain in the DPU Cluster Config state
If DPU objects remain in a DPU Cluster Config state and do not progress, the hosted cluster might have pending certificate signing requests (CSRs) that must be approved.
- Check for pending CSRs in the hosted cluster
- Switch to the hosted cluster context and check for any pending CSRs:
$ export KUBECONFIG=<path_to_hosted_cluster_kubeconfig>Review the output and approve any CSRs that show a$ oc get csr
Pendingstatus.
Management cluster nodes do not become ready
If management cluster nodes do not reach a Ready state after DPU provisioning completes, the OVN-Kubernetes CNI pods might not be running correctly on the management cluster or the hosted cluster.
- Check OVN-Kubernetes pods on the management cluster
- Switch to the management cluster context and verify that all OVN-Kubernetes pods are running on the x86_64 worker nodes and control plane nodes:
$ export KUBECONFIG=<path_to_management_cluster_kubeconfig>$ oc get pods -n openshift-ovn-kubernetes -o wide
- Check OVN-Kubernetes pods on the hosted cluster
- Switch to the hosted cluster context and verify that all OVN-Kubernetes pods are running on the DPU workers:
$ export KUBECONFIG=<path_to_hosted_cluster_kubeconfig>Ensure that all pods in the$ oc get pods -n openshift-ovn-kubernetes -o wide
openshift-ovn-kubernetesnamespace are in aRunningstate on both clusters.
DPU provisioning fails with BMC certificate errors
If DPU provisioning fails with certificate errors when you add worker nodes by using the Bare Metal Operator, the baseboard management controller (BMC) certificates might be untrusted or expired, or the BareMetalHost credentials might be incorrect.
- Verify BMC certificate validity
- Run the following command to inspect the BMC TLS certificate, replacing
<bmc_ip>with the BMC IP address and<bmc_hostname>with the BMC hostname:Update the certificates in the BMC configuration if they are expired or untrusted.$ openssl s_client -connect <bmc_ip>:443 -servername <bmc_hostname> - Verify BareMetalHost BMC credentials
- Ensure that the
BareMetalHostresource references the correct BMC secret and connection details, including the Redfish address and credentials for the worker server.
Worker node CSR approval fails
If certificate signing request (CSR) approval for worker nodes fails, network connectivity between the management cluster and the DPU or hosted cluster path might be incomplete.
- Check Host-Based Networking pods on worker nodes
- Run the following command to verify that HBN pods are running:
$ oc get pods -n openshift-hbn -o wide
- Verify DPU management network connectivity
- From a management cluster node, ping the DPU management IP address:
$ ping <dpu_management_ip>
- Verify the
br-exbridge on worker nodes - Confirm that the
br-exbridge that the workerMachineConfigresource creates is present and that required firewall rules allow traffic on the DPU management and high-speed networks.
DPU nodes remain NotReady in the hosted cluster
If DPU nodes remain in a NotReady state in the hosted cluster, DPU provisioning might be incomplete, or the DPU firmware and DOCA software versions might be incompatible with the DPF Operator version.
- Check DPU and DPU service status on the management cluster
- Run the following commands:
$ oc get dpu -n dpf-operator-system$ oc get dpuservice -n dpf-operator-system
- Verify node status in the hosted cluster
- Switch to the hosted cluster kubeconfig and list the nodes:
$ export KUBECONFIG=<path_to_hosted_cluster_kubeconfig>$ oc get nodes
- Check DPF Operator and related pod logs
- On the management cluster, inspect logs from DPF-related pods for provisioning or networking errors:
$ oc logs -n dpf-operator-system <dpu_related_pod_name>
- Verify firmware and software compatibility
- Confirm that the BlueField firmware and DOCA software versions on the DPU are compatible with the DPF Operator version that you deployed.
Troubleshoot hosted cluster issues
You can diagnose and resolve DPU hosted cluster issues, including CSR approval failures, kubeconfig access problems, and worker node join failures.
Prerequisites
- The DPF HCP Provisioner is installed and configured.
- DPU provisioning has completed on the management cluster.
- You have access to kubeconfig files for both the management cluster and the hosted cluster.
Procedure
-
Verify that the hosted cluster is accessible by running the following commands:
$ export KUBECONFIG=/path/to/hosted-cluster.kubeconfig$ oc cluster-infoIf the hosted cluster API server is not accessible, check the hosted control plane status on the management cluster.
-
Switch to the management cluster context and verify that the hosted control plane components are running:
$ export KUBECONFIG=/path/to/management-cluster.kubeconfig$ oc get pods -n clusters-$HOSTED_CLUSTER_NAMEVerify that the
etcd,kube-apiserver,kube-controller-manager, andkube-schedulerpods are all in aRunningstate. -
Check the DPF HCP Provisioner status for any error conditions:
$ oc get dpfhcpprovisioner -n dpf-operator-system -o yamlReview the
status.conditionsfield for any conditions that indicate a failure. -
Return to the hosted cluster context and check for pending CSRs:
$ export KUBECONFIG=/path/to/hosted-cluster.kubeconfig$ oc get csr --sort-by=.metadata.creationTimestampExample outputNAME AGE SIGNERNAME REQUESTOR CONDITIONcsr-abc12 30s kubernetes.io/kubelet-serving system:node:dpu-worker1 Pendingcsr-def34 25s kubernetes.io/kube-apiserver-client-kubelet system:bootstrap:abc123 Pending -
Approve any pending CSRs. To approve a single CSR, run the following command, replacing
<csr-name>with the CSR name:$ oc adm certificate approve <csr-name>To approve all pending CSRs in a single command, run:
$ oc get csr -o name | xargs oc adm certificate approve -
Verify that the DPU workers are joining the hosted cluster:
$ oc get nodesExample outputNAME STATUS ROLES AGE VERSIONdpu-worker1 Ready worker 5m v1.35.6dpu-worker2 Ready worker 5m v1.35.6 -
If nodes are not joining, check whether the bootstrap token is still valid by running the following command on the hosted cluster:
$ oc get secrets -n kube-system | grep bootstrap-tokenBootstrap tokens have a limited lifetime. The DPF HCP Provisioner should create new tokens automatically. If tokens are expired and not being renewed, check the provisioner logs for errors.
-
Check the kubelet logs on the DPU for authentication or certificate errors. Switch to the management cluster context and open a debug shell on the DPU-enabled worker node:
$ export KUBECONFIG=/path/to/management-cluster.kubeconfig$ oc debug node/<dpu-enabled-worker-node>Inside the debug shell, run the following commands to stream the kubelet logs:
sh-5.1# chroot /hostsh-5.1# journalctl -u kubelet -fLook for authentication errors or certificate-related failures in the log output.
-
Verify that OVN-Kubernetes is running correctly on the hosted cluster. Switch to the hosted cluster context and run the following command:
$ export KUBECONFIG=/path/to/hosted-cluster.kubeconfig$ oc get pods -n openshift-ovn-kubernetes -o wideEnsure that OVN-Kubernetes pods are running on the DPU ARM cores.
Troubleshooting
CSR approval failures: Verify that the DPF HCP Provisioner has the required RBAC permissions to approve CSRs, check the provisioner logs for certificate-related errors, and ensure that the cluster CA is configured correctly.
Node join failures: Verify that the bootstrap kubeconfig was correctly generated by the provisioner, check network connectivity between the DPUs and the hosted control plane API server, and ensure that the kubelet configuration includes the correct API server endpoint.
Control plane access issues: Verify that the hosted cluster virtual IP address is configured and accessible, check the LoadBalancer service status for the hosted API server, and ensure that MetalLB is correctly configured and announcing the VIP.
Network connectivity problems: Verify the VTEP network configuration between DPUs, check that the DPU high-speed network interfaces are operational, and ensure that the required ports are open for inter-DPU communication.
Troubleshoot DPF networking issues
You can diagnose and resolve DPF networking issues, including OVN-Kubernetes configuration problems, MTU mismatches, and connectivity failures.
Prerequisites
- DPU provisioning completed successfully.
- The hosted cluster is accessible with DPU worker nodes joined.
- You have access to both management and hosted cluster contexts.
Procedure
-
Verify OVN-Kubernetes pod status on the management cluster:
$ export KUBECONFIG=/path/to/management-cluster.kubeconfig$ oc get pods -n openshift-ovn-kubernetes -o wideCheck that the following pods are running:
ovnkube-control-plane-*pods are running on control plane nodes only.ovnkube-node-*pods are running on all nodes.ovs-node-*pods are running on all nodes.
-
Check OVN-Kubernetes configuration on the hosted cluster:
$ export KUBECONFIG=/path/to/hosted-cluster.kubeconfig$ oc get pods -n openshift-ovn-kubernetes -o wideVerify that OVN-Kubernetes pods are running on DPU ARM cores, not on host x86 CPUs.
-
Verify the network MTU configuration:
$ oc get network.operator.openshift.io cluster -o yaml | grep -A 5 defaultNetworkCheck the following MTU values:
- Standard networks: MTU 1400 for pods, 1500 for nodes.
- Jumbo frame networks: MTU 8940 for pods, 9000 for nodes.
-
Test basic pod-to-pod connectivity:
$ export KUBECONFIG=/path/to/hosted-cluster.kubeconfig$ oc run test-pod-1 --image=nicolaka/netshoot --rm -it -- /bin/bashFrom another terminal, run:
$ oc run test-pod-2 --image=nicolaka/netshoot --rm -it -- /bin/bashTest connectivity between the pods by using cluster IP addresses.
-
Check VTEP network configuration:
$ export KUBECONFIG=/path/to/management-cluster.kubeconfig$ oc debug node/<dpu-enabled-worker-node>In the debug shell, run:
$ chroot /host$ ip addr show | grep $VTEP_CIDRVerify that VTEP interfaces are configured with the correct IP addresses from the
VTEP_CIDRrange. -
Test VTEP connectivity:
$ ping -c 4 <other-dpu-vtep-ip>If the ping fails, check routing and firewall rules between DPU nodes.
-
Verify OVN database connectivity:
$ export KUBECONFIG=/path/to/hosted-cluster.kubeconfig$ oc exec -n openshift-ovn-kubernetes <ovnkube-node-pod> -- ovn-nbctl showThe output should display the OVN logical network topology.
-
Check OVN-Kubernetes log errors:
$ oc logs -n openshift-ovn-kubernetes <ovnkube-node-pod> -c ovn-controllerLook for the following error types:
- Database connectivity issues
- Port binding failures
- Flow programming errors
-
Verify service mesh connectivity:
$ export KUBECONFIG=/path/to/hosted-cluster.kubeconfig$ oc create service clusterip test-svc --tcp=80:80$ oc run test-client --image=nicolaka/netshoot --rm -it -- nc -vz test-svc 80A successful connection indicates that service traffic is flowing through the DPU data plane.
-
Check the SR-IOV network device plugin:
$ export KUBECONFIG=/path/to/management-cluster.kubeconfig$ oc get sriovnetworknodepolicy -n openshift-sriov-network-operatorVerify that SR-IOV policies are correctly applied to DPU-enabled worker nodes.
Troubleshooting
OVN-Kubernetes pod failures
Check that the OVN Helm chart version is compatible with your OpenShift Container Platform version. Verify that the CNI configuration matches the DPU acceleration requirements. Ensure that OVN databases are accessible from the DPU worker nodes.
MTU mismatch issues
Verify that all network components use consistent MTU values. Check that the physical network infrastructure supports the configured MTU. Update MTU values if the network environment has changed.
VTEP connectivity problems
Verify that the VTEP CIDR does not conflict with existing network ranges. Check that routing is configured between DPU nodes. Ensure that firewalls allow VTEP traffic on the required ports.
Service connectivity failures
Verify that kube-proxy is correctly configured on DPU nodes. Check that iptables rules are correctly programmed. Ensure that DPU acceleration is correctly handling service traffic.
SR-IOV configuration issues
Verify that the SR-IOV Operator is compatible with the DPU firmware. Check that the VF count matches the configured value. Ensure that VFs are correctly allocated to the correct namespaces.
DPF diagnostic commands and log collection
You can run diagnostic commands and collect logs to investigate DPF component status, DPU provisioning failures, and hosted cluster issues when troubleshooting or opening support cases.
Quick status overview commands
The following commands provide a quick overview of DPF component status:
$ oc get dpudeployment,dpuservicetemplate,dpuserviceconfiguration,bfb,dpu -n dpf-operator-system
$ oc get pods -n dpf-operator-system -l app.kubernetes.io/part-of=dpf-operator
$ oc get nodes -l feature.node.kubernetes.io/dpu-enabled="" \
-o custom-columns=NAME:.metadata.name,STATUS:.status.conditions[?(@.type=="Ready")].status,AGE:.metadata.creationTimestamp
$ oc get hostedcluster -n clusters-$HOSTED_CLUSTER_NAME
$ oc get nodepool -n clusters-$HOSTED_CLUSTER_NAME
Detailed diagnostic commands
$ oc describe dpu -n dpf-operator-system
$ oc describe bfb -n dpf-operator-system bf-bundle
$ oc get dpfoperatorconfig -n dpf-operator-system -o yaml
$ oc describe dpuservicetemplate -n dpf-operator-system
$ oc describe dpuserviceconfiguration -n dpf-operator-system
$ oc get dpuservice -n dpf-operator-system -o wide
$ oc describe dpuservice -n dpf-operator-system
$ oc get nodefeaturerule -n openshift-nfd
$ oc describe node <worker-node> | grep -A 20 "Labels:"
Log collection commands
$ oc logs -n dpf-operator-system -l app.kubernetes.io/name=dpf-operator --tail=200 > dpf-operator.log
$ oc logs -n dpf-operator-system -l app.kubernetes.io/name=dpfhcp-provisioner-operator --tail=200 > dpfhcp-provisioner.log
$ oc debug node/<dpu-worker-node>
In the debug shell, run:
$ chroot /host
$ journalctl -u kubelet --since "1 hour ago" > kubelet.log
$ oc logs -n openshift-ovn-kubernetes -l app=ovnkube-node --tail=100 > ovn-kubernetes.log
$ oc logs -n openshift-sriov-network-operator -l app=sriov-network-operator --tail=100 > sriov-operator.log
$ oc logs -n openshift-nfd -l app=nfd-worker --tail=100 > nfd.log
System information collection
$ oc debug node/<dpu-worker-node>
In the debug shell, run:
$ lspci | grep -i mellanox
$ lshw -class network
$ dmidecode -t system
$ mlxfwmanager --query
$ mst status
$ ip addr show
$ ip route show
$ ethtool -i <interface>
Performance monitoring commands
$ oc exec -n dpf-operator-system <dts-service-pod> -- \
curl -s localhost:9189/metrics | grep -E "(current_link_speed|p[01]_eth_)"
$ oc adm top pods -n dpf-operator-system --containers
$ oc adm top nodes -l feature.node.kubernetes.io/dpu-enabled=""
Support information package
When opening a support case, collect the following information:
Environment information
- OpenShift Container Platform cluster version and build
- DPF Operator version and configuration
- Hardware specifications (server model, DPU model, firmware versions)
- Network topology and configuration
Configuration files
- DPF Operator configuration (
dpfoperatorconfig) - Service templates and configurations
- Network policies and configurations
- Environment variables used during installation
Log files
- DPF Operator logs (past 24 hours)
- Worker node system logs (past 4 hours)
- Kubernetes event logs related to DPF resources
- Application logs for affected services
Common log analysis patterns
Look for the following patterns in logs when troubleshooting:
DPU provisioning issues:
Error downloading BFB imageFailed to detect DPU hardwareProvisioning timeout exceeded
Networking issues:
OVN database connection failedFailed to program flowsInterface binding failed
Service deployment issues:
Image pull failedInsufficient resourcesConfigMap not found
Authentication issues:
Certificate signing request deniedUnauthorized access to API serverToken validation failed
Additional resources