Skip to content

NVIDIA DPF Operator release notes

Use the release notes to learn what is new or changed in the NVIDIA DPF Operator.

Release notes for NVIDIA DPF Operator 26.4.1

The NVIDIA DPF Operator on OpenShift Container Platform has known limitations for uninstall, secrets, MTU, secure boot, multi-DPU hosts, and HBN, unsupported OVN-Kubernetes features, and issues that can affect Grafana and DTS metrics.

DPF Operator v26.4.1

New features and enhancements

Enhanced observability with DTS integration
Added comprehensive DPU telemetry monitoring through the DOCA Telemetry Service (DTS) with built-in OpenShift Container Platform Console dashboard integration. DTS metrics are now accessible directly through the OpenShift Container Platform web console without requiring additional tools.
Improved hosted control planes integration
The DPF HCP Provisioner Operator provides enhanced lifecycle management for DPU hosted clusters, including automatic CSR approval, kubeconfig injection, and BlueField container image lookup.
Advanced traffic validation
A comprehensive traffic validation framework with pre-configured test pods uses nicolaka/netshoot containers to validate end-to-end DPU service chain functionality.
Enhanced troubleshooting capabilities
Expanded diagnostic tools and troubleshooting procedures cover DPU provisioning, hosted cluster management, networking issues, and comprehensive log collection.

Bug fixes

Improved BFB image handling
Fixed issues with BlueField Bootstream File (BFB) image download and verification processes.
Enhanced worker node detection
Resolved Node Feature Discovery (NFD) compatibility issues for reliable DPU hardware detection.
Networking stability improvements
Fixed OVN-Kubernetes integration issues that could cause worker nodes to remain in NotReady state.

Known issues and limitations

Only x86_64 workers are supported
Only x86_64 worker nodes are supported in this release. ARM-based DPU workers are not supported.
DPF Operator uninstall is not supported
The DPF Operator does not support an automated uninstall. If you must remove DPF, set spec.manageDPUServiceTemplates to false in the DPFHCPProvisionerConfig resource before you uninstall. This prevents the DPF HCP Provisioner Operator from continuing to manage DPUServiceTemplate resources during the uninstall process.
Secret references are immutable
The pull secret and SSH secret references are immutable after creation and cannot be modified. Ensure that each secret contains the correct data before you create it and reference it in the DPFHCPProvisioner custom resource.
Secondary pod interfaces are not supported
Secondary pod interfaces (MultiNetwork) are not supported.
MTU changes are not supported after deployment
You cannot change the MTU value after deployment.
controlPlaneMTU and highSpeedMTU must use the same value
In the DPFOperatorConfig custom resource, you must set controlPlaneMTU and highSpeedMTU to the same value, either 1500 or 9000.
Secure boot firmware requirement
To boot the RHCOS BFB image with secure boot enabled, the DPU firmware must be at version 3.1.0 or later. Use a BFB firmware bundle to upgrade the firmware.
Multi-DPU hosts are not supported
Hosts with more than one DPU are not supported.
Redeploying a DPUDeployment is not supported
Redeploying a DPUDeployment is not supported in this release.
Deployments cannot target all nodes in a cluster
Because of a limitation in the resource injector, a deployment can target either the DPU workers or all other nodes, but not both.
Host-Based Networking (HBN) pods stuck in FailedCreatePodSandBox
An HBN daemon set pod might remain in the FailedCreatePodSandBox state. As a workaround, delete and re-create the affected pods. For more information, see OCPBUGS-100251.
HostedCluster upgrade during an in-progress upgrade
Changing the ocpReleaseImage of a HostedCluster while an upgrade is already in progress is not supported.
Connectivity loss after a DPU reboot or upgrade

When a DPU reboots, the corresponding DPU worker node loses connectivity and a NoExecute taint is added to the host. Most pods are evicted immediately, but some daemon set pods remain and might lose connectivity until you re-create them. For example:

Example output
openshift-network-diagnostics   network-check-target-lpkp2   0/1   Running
DPU stuck in the NodeEffect or Initializing state

The NVIDIA Maintenance Operator might fail to pause the machine config pool, which leaves the DPU in the NodeEffect or Initializing state. As a workaround, pause the worker-dpu machine config pool manually:

$ oc patch mcp worker-dpu --type merge -p '{
  "spec": {"paused": true},
  "metadata": {"annotations": {"maintenance.nvidia.com/mcp-paused": "true"}}
}'
Workload pods do not recover after an IPMI reset reboot
After an IPMI reset reboot, workload pods might fail to recover because of a known kubelet bug (Kubernetes issue 128043) that prevents virtual function (VF) devices from being re-created immediately at startup. This does not break functionality, but it leaves the cluster in an inconsistent state. Standard and IPMI2 reboots recover cleanly. As a workaround, re-create the affected pods manually if needed.
The SR-IOV device plugin can report fewer virtual functions than configured
After a node reboot or DPU redeployment, the SR-IOV device plugin might publish the node’s virtual function (VF) resource count before all VFs are created. The init container unblocks when the first VF appears instead of waiting for all configured VFs, so the reported openshift.io/bf3_vfs capacity can be lower than expected. As a workaround, restart the SR-IOV device plugin pod on the affected node, after which the full count is reported.

OVN-Kubernetes feature support

The following table lists the support and hardware offload status of OVN-Kubernetes features in this release.

OVN-Kubernetes feature support and offload status

Feature Supported Offloaded
Administrative Network Policies (ANP) Yes Yes
Egress IP Yes No
Egress Firewall Yes No
Egress Quality of Service (QoS) Yes No
Secondary networks No No
User Defined Networks (UDN) No No
Quality of Service (QoS) No No
Multiple External Gateways (MEG) No No
OVN-Kubernetes identity No No
Border Gateway Protocol (BGP) No No
Multicast No No
Hybrid Overlay No No
Local gateway mode No No
IPFIX or NetFlow sampling No No

Grafana deployment issues

Grafana shows that the application is not available

Grafana pods might be scheduled on worker nodes that depend on DPU networking, which creates a circular dependency.

Configure Grafana to run on control plane nodes by adding nodeSelector and tolerations to the Grafana custom resource:

spec:
  deployment:
    spec:
      template:
        spec:
          nodeSelector:
            node-role.kubernetes.io/control-plane: ""
          tolerations:
          - key: node-role.kubernetes.io/master
            operator: Exists
            effect: NoSchedule
          - key: node-role.kubernetes.io/control-plane
            operator: Exists
            effect: NoSchedule

DTS metrics collection issues

DTS metrics are not appearing in Prometheus or Grafana

The ServiceMonitor might not be configured correctly, or DTS pods might not be running.

Verify that DTS pods are running:

$ oc get pods -n dpf-operator-system -l app=dts

Check the ServiceMonitor configuration:

$ oc get servicemonitor -n dpf-operator-system
$ oc describe servicemonitor <servicemonitor-name> -n dpf-operator-system

Verify that user workload monitoring is enabled:

$ oc get configmap cluster-monitoring-config -n openshift-monitoring -o yaml

Check Prometheus targets to ensure that DTS endpoints are being scraped. Access the Prometheus web console and go to Status → Targets to verify that DTS endpoints are listed and healthy.