NVIDIA DPF Operator release notes¶
Use the release notes to learn what is new or changed in the NVIDIA DPF Operator.
Release notes for NVIDIA DPF Operator 26.4.1¶
The NVIDIA DPF Operator on OpenShift Container Platform has known limitations for uninstall, secrets, MTU, secure boot, multi-DPU hosts, and HBN, unsupported OVN-Kubernetes features, and issues that can affect Grafana and DTS metrics.
DPF Operator v26.4.1¶
New features and enhancements
- Enhanced observability with DTS integration
- Added comprehensive DPU telemetry monitoring through the DOCA Telemetry Service (DTS) with built-in OpenShift Container Platform Console dashboard integration. DTS metrics are now accessible directly through the OpenShift Container Platform web console without requiring additional tools.
- Improved hosted control planes integration
- The DPF HCP Provisioner Operator provides enhanced lifecycle management for DPU hosted clusters, including automatic CSR approval, kubeconfig injection, and BlueField container image lookup.
- Advanced traffic validation
- A comprehensive traffic validation framework with pre-configured test pods uses
nicolaka/netshootcontainers to validate end-to-end DPU service chain functionality. - Enhanced troubleshooting capabilities
- Expanded diagnostic tools and troubleshooting procedures cover DPU provisioning, hosted cluster management, networking issues, and comprehensive log collection.
Bug fixes
- Improved BFB image handling
- Fixed issues with BlueField Bootstream File (BFB) image download and verification processes.
- Enhanced worker node detection
- Resolved Node Feature Discovery (NFD) compatibility issues for reliable DPU hardware detection.
- Networking stability improvements
- Fixed OVN-Kubernetes integration issues that could cause worker nodes to remain in
NotReadystate.
Known issues and limitations
- Only x86_64 workers are supported
- Only x86_64 worker nodes are supported in this release. ARM-based DPU workers are not supported.
- DPF Operator uninstall is not supported
- The DPF Operator does not support an automated uninstall. If you must remove DPF, set
spec.manageDPUServiceTemplatestofalsein theDPFHCPProvisionerConfigresource before you uninstall. This prevents the DPF HCP Provisioner Operator from continuing to manageDPUServiceTemplateresources during the uninstall process. - Secret references are immutable
- The pull secret and SSH secret references are immutable after creation and cannot be modified. Ensure that each secret contains the correct data before you create it and reference it in the
DPFHCPProvisionercustom resource. - Secondary pod interfaces are not supported
- Secondary pod interfaces (MultiNetwork) are not supported.
- MTU changes are not supported after deployment
- You cannot change the MTU value after deployment.
controlPlaneMTUandhighSpeedMTUmust use the same value- In the
DPFOperatorConfigcustom resource, you must setcontrolPlaneMTUandhighSpeedMTUto the same value, either1500or9000. - Secure boot firmware requirement
- To boot the RHCOS BFB image with secure boot enabled, the DPU firmware must be at version 3.1.0 or later. Use a BFB firmware bundle to upgrade the firmware.
- Multi-DPU hosts are not supported
- Hosts with more than one DPU are not supported.
- Redeploying a DPUDeployment is not supported
- Redeploying a
DPUDeploymentis not supported in this release. - Deployments cannot target all nodes in a cluster
- Because of a limitation in the resource injector, a deployment can target either the DPU workers or all other nodes, but not both.
- Host-Based Networking (HBN) pods stuck in FailedCreatePodSandBox
- An HBN daemon set pod might remain in the
FailedCreatePodSandBoxstate. As a workaround, delete and re-create the affected pods. For more information, see OCPBUGS-100251. - HostedCluster upgrade during an in-progress upgrade
- Changing the
ocpReleaseImageof aHostedClusterwhile an upgrade is already in progress is not supported. - Connectivity loss after a DPU reboot or upgrade
-
When a DPU reboots, the corresponding DPU worker node loses connectivity and a
NoExecutetaint is added to the host. Most pods are evicted immediately, but some daemon set pods remain and might lose connectivity until you re-create them. For example: - DPU stuck in the NodeEffect or Initializing state
-
The NVIDIA Maintenance Operator might fail to pause the machine config pool, which leaves the DPU in the
NodeEffectorInitializingstate. As a workaround, pause theworker-dpumachine config pool manually: - Workload pods do not recover after an IPMI reset reboot
- After an IPMI reset reboot, workload pods might fail to recover because of a known kubelet bug (Kubernetes issue 128043) that prevents virtual function (VF) devices from being re-created immediately at startup. This does not break functionality, but it leaves the cluster in an inconsistent state. Standard and IPMI2 reboots recover cleanly. As a workaround, re-create the affected pods manually if needed.
- The SR-IOV device plugin can report fewer virtual functions than configured
- After a node reboot or DPU redeployment, the SR-IOV device plugin might publish the node’s virtual function (VF) resource count before all VFs are created. The init container unblocks when the first VF appears instead of waiting for all configured VFs, so the reported
openshift.io/bf3_vfscapacity can be lower than expected. As a workaround, restart the SR-IOV device plugin pod on the affected node, after which the full count is reported.
OVN-Kubernetes feature support
The following table lists the support and hardware offload status of OVN-Kubernetes features in this release.
OVN-Kubernetes feature support and offload status
| Feature | Supported | Offloaded |
|---|---|---|
| Administrative Network Policies (ANP) | Yes | Yes |
| Egress IP | Yes | No |
| Egress Firewall | Yes | No |
| Egress Quality of Service (QoS) | Yes | No |
| Secondary networks | No | No |
| User Defined Networks (UDN) | No | No |
| Quality of Service (QoS) | No | No |
| Multiple External Gateways (MEG) | No | No |
| OVN-Kubernetes identity | No | No |
| Border Gateway Protocol (BGP) | No | No |
| Multicast | No | No |
| Hybrid Overlay | No | No |
| Local gateway mode | No | No |
| IPFIX or NetFlow sampling | No | No |
Grafana deployment issues
- Grafana shows that the application is not available
-
Grafana pods might be scheduled on worker nodes that depend on DPU networking, which creates a circular dependency.
Configure Grafana to run on control plane nodes by adding
nodeSelectorandtolerationsto the Grafana custom resource:
DTS metrics collection issues
- DTS metrics are not appearing in Prometheus or Grafana
-
The
ServiceMonitormight not be configured correctly, or DTS pods might not be running.Verify that DTS pods are running:
Check the
ServiceMonitorconfiguration:$ oc get servicemonitor -n dpf-operator-system $ oc describe servicemonitor <servicemonitor-name> -n dpf-operator-systemVerify that user workload monitoring is enabled:
Check Prometheus targets to ensure that DTS endpoints are being scraped. Access the Prometheus web console and go to Status → Targets to verify that DTS endpoints are listed and healthy.