Set up the environment for DPF¶
Before installing the NVIDIA DPF Operator, you must set up the management cluster, configure worker nodes, and install and configure the required Operators.
Set up the management cluster¶
The management cluster is a standard OpenShift Container Platform 4.22 cluster installed by using the Assisted Installer. This cluster hosts the DPF Operators and the hosted control planes for managing the hosted cluster on DPUs.
Prerequisites
- You have access to the Red Hat Hybrid Cloud Console.
- You have the OpenShift CLI (
oc) installed.
Procedure
-
Go to the Red Hat Hybrid Cloud Console cluster creation page and create a cluster with control-plane nodes only. Select Data center → Assisted Installer.
-
Optional: Configure jumbo MTU for each control plane node.
-
Under Hosts' network configuration in the Assisted Installer wizard, select Static IP, bridges, and bonds.
-
Set the Static network configurations section per node according to the following template, using the relevant MAC address and interface name for each node:
interfaces: - ipv4: dhcp: true enabled: true mac-address: <xx:xx:xx:xx:xx:xx> mtu: 1500 # Set to 1500 for standard MTU or 9000 for jumbo frames name: <interface-name> state: up type: ethernetNote
- You can alternatively configure MTU allocation on the DHCP server that allocates IPs to the control plane nodes.
- If virtual machines are used for control-plane nodes, the MTU must be set on the bridge of the hypervisor used by the VMs.
- When using MTU 9000, ensure the switch ports that connect the cluster’s control-plane nodes are set to handle jumbo frames.
-
-
Select the following operators to install with the cluster:
- Storage → Logical Volume Manager Storage
- Platform Operations & Lifecycle → MultiCluster Engine
- Scheduling → Node Feature Discovery
-
Click Add hosts to add hosts to the cluster. Only control plane nodes are required at this stage.
-
After the installation completes, download the
KUBECONFIGfile and save it asmgmt-kubeconfig.
Verification
-
Set the
KUBECONFIGenvironment variable: -
Verify that all nodes are in a
Readystate:
Configure worker nodes for DPU operation¶
You must deploy the dpu-worker-config Helm chart to configure worker nodes with DPUs before adding those nodes to the management cluster. The dpu-worker-config Helm chart creates the MachineConfigPool, dpu-worker-configuration MachineConfig, and other required resources that configure the bridge, OVS services, and IP routing on worker nodes. The MachineConfigPool groups DPU-equipped worker nodes so that the Machine Config Operator can apply DPU-specific configurations to them.
The MachineConfig resource performs several configuration tasks required by DPF:
- Bridge Configuration: Creates a
br-exbridge interface that enables communication between the DPU and the hosted cluster control plane running on the management cluster. This name must match thedpuNodeOOBBridgeNamevalue in theDPFOperatorConfigresource, or DPU provisioning fails. For more information refer to DPF Operator Prerequisites. - OVS Service Management: Disables OpenShift’s default OVS services on x86 worker nodes. This is required for OVN-Kubernetes DPU Host mode operation, where networking functions are offloaded to the DPU rather than running on the host CPU.
- IP Rules Configuration: Sets routing rules required for pod-to-host control-plane traffic.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - You have installed the OpenShift CLI (
oc). - The Helm CLI (
helm) is installed on your workstation. - You have a pull secret file that includes credentials for
registry.redhat.io. Helm reads its registry credentials from this file, which is separate from the container runtime configuration.
Procedure
-
Set the
OPENSHIFT_PULL_SECRETenvironment variable to the path of your pull secret file: -
Deploy the
dpu-worker-configHelm chart:
Verification
-
Verify that the
dpu-worker-configHelm release is deployed: -
Verify that the
MachineConfigPoolwas created automatically: -
Verify that the
dpu-worker-configurationMachineConfig is created:Note
The Machine Config Operator automatically reboots worker nodes to apply DPU-specific configurations after nodes with the
worker-dpulabel are added to the cluster.
Create the DPF namespace¶
You must create a dedicated namespace for the DPF Operator and its components before installing the Operator.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - You have installed the OpenShift CLI (
oc).
Procedure
-
Create the
dpf-operator-systemnamespace:
Verification
-
Verify that the namespace was created:
Required Operators¶
Before you install the DPF Operator, you must install the cert-manager Operator for Red Hat OpenShift, MetalLB Operator, Red Hat OpenShift GitOps, and NVIDIA Maintenance Operator.
The multicluster engine Operator and the Node Feature Discovery Operator can be installed during management cluster creation by using the Assisted Installer. If you did not install the multicluster engine Operator, follow "Install the multicluster engine Operator" before configuring the required Operators.
Install the multicluster engine Operator¶
If you did not install the multicluster engine Operator by using the Assisted Installer, install it from the OpenShift CLI. You do not need to install Red Hat Advanced Cluster Management or create a MultiClusterHub resource to use the multicluster engine Operator. If the Assisted Installer already created a MultiClusterEngine resource, skip this procedure.
Prerequisites
- You have access to the management cluster as a user with the
cluster-adminrole. - You have installed the OpenShift CLI (
oc). - No
MultiClusterEngineresource exists on the cluster. - If you already have an existing
MultiClusterEngineresource with a different name, note that the following verification commands usemce. Substitute your resource’s name where needed.
Procedure
-
Check which multicluster engine channels are available from the
redhat-operatorscatalog:$ oc get packagemanifests -n openshift-marketplace -l catalog=redhat-operators \ -o jsonpath='{.items[?(@.metadata.name=="multicluster-engine")].status.channels[*].name}{"\n"}'The following example uses
stable-2.17, which is the channel used by this deployment. If your catalog does not offer this channel, select a supported channel for your OpenShift version before continuing. -
Create a file named
mce-operator.yamlwith the following content:apiVersion: v1 kind: Namespace metadata: name: multicluster-engine --- apiVersion: operators.coreos.com/v1 kind: OperatorGroup metadata: name: multicluster-engine namespace: multicluster-engine spec: targetNamespaces: - multicluster-engine --- apiVersion: operators.coreos.com/v1alpha1 kind: Subscription metadata: name: multicluster-engine namespace: multicluster-engine spec: channel: stable-2.17 name: multicluster-engine source: redhat-operators sourceNamespace: openshift-marketplace installPlanApproval: AutomaticIf you selected a different supported channel, replace
stable-2.17with that channel in the file. -
Apply the file:
-
Wait until the multicluster engine Operator CSV reports
Succeededand theMultiClusterEngineCRD is established:$ oc get csv -n multicluster-engine $ oc wait crd/multiclusterengines.multicluster.openshift.io --for=create --timeout=10m $ oc wait crd/multiclusterengines.multicluster.openshift.io --for=condition=Established --timeout=5mRepeat the CSV check until its phase is
Succeededbefore creating the custom resource. -
Create a file named
mce.yamlwith the following content: -
Apply the file:
Verification
-
Wait for the
MultiClusterEngineresource to become available: -
Verify that the hosted control plane component is enabled:
Additional resources
Install the cert-manager Operator¶
The cert-manager Operator for Red Hat OpenShift manages TLS certificates for DPF components. You install this operator by using the OpenShift CLI.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - You have installed the OpenShift CLI (
oc).
Procedure
-
Create a file named
cert-manager-operator.yamlwith the following content:apiVersion: v1 kind: Namespace metadata: name: cert-manager --- apiVersion: operators.coreos.com/v1 kind: OperatorGroup metadata: name: openshift-cert-manager-operator namespace: cert-manager spec: targetNamespaces: - cert-manager --- apiVersion: operators.coreos.com/v1alpha1 kind: Subscription metadata: name: openshift-cert-manager-operator namespace: cert-manager spec: channel: stable-v1 name: openshift-cert-manager-operator source: redhat-operators sourceNamespace: openshift-marketplace -
Apply the file:
Verification
-
Verify that the Operator is installed:
Install the MetalLB Operator¶
The MetalLB Operator provides load balancing services for DPF components on the management cluster. You install this operator by using the OpenShift CLI.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - You have installed the OpenShift CLI (
oc).
Procedure
-
Create a file named
metallb-operator.yamlwith the following content:apiVersion: v1 kind: Namespace metadata: name: metallb-system --- apiVersion: operators.coreos.com/v1alpha1 kind: Subscription metadata: name: metallb-operator namespace: openshift-operators spec: channel: "stable" name: metallb-operator source: redhat-operators sourceNamespace: openshift-marketplace installPlanApproval: Automatic config: # Tolerate the taint on the master nodes tolerations: - key: "node-role.kubernetes.io/control-plane" operator: "Exists" effect: "NoSchedule" # Force scheduling only on nodes with the control-plane label affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: "node-role.kubernetes.io/control-plane" operator: "Exists" -
Apply the file:
Verification
-
Wait for the MetalLB custom resource definition to be created and established before configuring MetalLB:
Install the GitOps Operator¶
The Red Hat OpenShift GitOps manages DPF service deployments and configurations using GitOps principles. You install this operator by using the OpenShift CLI.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - You have installed the OpenShift CLI (
oc).
Procedure
-
Create a file named
gitops-operator.yamlwith the following content:apiVersion: v1 kind: Namespace metadata: name: openshift-gitops-operator labels: openshift.io/cluster-monitoring: "true" --- apiVersion: operators.coreos.com/v1 kind: OperatorGroup metadata: name: openshift-gitops-operator namespace: openshift-gitops-operator spec: upgradeStrategy: Default --- apiVersion: operators.coreos.com/v1alpha1 kind: Subscription metadata: name: openshift-gitops-operator namespace: openshift-gitops-operator spec: channel: gitops-1.21 config: env: - name: ARGOCD_CLUSTER_CONFIG_NAMESPACES value: "openshift-gitops,dpf-operator-system" - name: CONTROLLER_CLUSTER_ROLE value: "cluster-admin" - name: SERVER_CLUSTER_ROLE value: "cluster-admin" installPlanApproval: Automatic name: openshift-gitops-operator source: redhat-operators sourceNamespace: openshift-marketplace -
Apply the file:
Verification
-
Wait for the Argo CD custom resource definition to be created and established before creating an
ArgoCDresource:
Install the NVIDIA Maintenance Operator¶
The NVIDIA Maintenance Operator assists in performing maintenance tasks and gracefully draining DPU worker nodes. You install this operator by using Helm.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - You have installed the OpenShift CLI (
oc). - You have installed the Helm CLI (
helm).
Procedure
-
Create a Helm values file named
maintenance-operator-values.yamlwith the following content:operatorConfig: maxParallelOperations: 60% operator: affinity: nodeAffinity: requiredDuringSchedulingIgnoredDuringExecution: nodeSelectorTerms: - matchExpressions: - key: "node-role.kubernetes.io/master" operator: Exists - matchExpressions: - key: "node-role.kubernetes.io/control-plane" operator: Exists tolerations: - key: node-role.kubernetes.io/master operator: Exists effect: NoSchedule - key: node-role.kubernetes.io/control-plane operator: Exists effect: NoSchedule -
Install the Operator by using Helm:
Verification
-
Verify that the Operator pod is running:
Additional resources
Configure the required Operators¶
After the required Operators are installed, configure Node Feature Discovery, MetalLB, GitOps, and Cluster Network Operator for the DPF environment. This procedure also verifies that the multicluster engine and hosted control plane components are ready.
Prerequisites
- You have access to the cluster as a user with the
cluster-adminrole. - You have installed the OpenShift CLI (
oc). - You have installed the cert-manager Operator for Red Hat OpenShift, MetalLB Operator, Red Hat OpenShift GitOps, and NVIDIA Maintenance Operator.
- You have installed the Logical Volume Manager Storage Operator, multicluster engine Operator, and the Node Feature Discovery Operator. You can install them by using the Assisted Installer during cluster creation. If you did not install the multicluster engine Operator, follow "Install the multicluster engine Operator" before continuing.
Procedure
-
Define the cluster variables:
$ export CLUSTER_NAME="doca-mgmt" $ export BASE_DOMAIN="example.com" $ export HOST_CLUSTER_API="api.${CLUSTER_NAME}.${BASE_DOMAIN}"where:
CLUSTER_NAME- Specifies the management cluster name.
BASE_DOMAIN- Specifies the management cluster base domain.
HOST_CLUSTER_API- Specifies the management cluster API endpoint.
-
Create a file named
nfd-instance.yamlwith the followingNodeFeatureDiscoveryresource definition:apiVersion: nfd.openshift.io/v1 kind: NodeFeatureDiscovery metadata: name: nfd-instance namespace: openshift-nfd spec: operand: workerEnvs: - name: KUBERNETES_SERVICE_HOST value: $HOST_CLUSTER_API - name: KUBERNETES_SERVICE_PORT value: "6443" workerConfig: configData: | sources: pci: deviceClassWhitelist: - "0200" - "03" - "12" - "0207" deviceLabelFields: - "vendor" - "device" - "class" -
Apply the file by using
envsubstto substitute the environment variables: -
Create a file named
nfd-rule.yamlwith the followingNodeFeatureRuleresource definition:apiVersion: nfd.openshift.io/v1alpha1 kind: NodeFeatureRule metadata: name: dpu-detection-rule namespace: openshift-nfd spec: rules: - labels: dpu-enabled: "" matchFeatures: - feature: pci.device matchExpressions: device: op: In value: - a2d6 - a2dc vendor: op: In value: - 15b3 name: DPU-detection-rule -
Apply the file:
-
Create a file named
metallb-config.yamlwith the followingMetalLBresource definition: -
Apply the MetalLB resource file:
-
Create a file named
argocd-instance.yamlwith the followingArgoCDresource definition:apiVersion: argoproj.io/v1beta1 kind: ArgoCD metadata: name: argocd namespace: dpf-operator-system spec: nodePlacement: nodeSelector: node-role.kubernetes.io/control-plane: "" tolerations: - key: node-role.kubernetes.io/master operator: Exists effect: NoSchedule - key: node-role.kubernetes.io/control-plane operator: Exists effect: NoSchedule server: route: enabled: true controller: {} repo: {} applicationSet: enabled: false resourceExclusions: | - apiGroups: - packages.operators.coreos.com kinds: - PackageManifest sso: provider: dex dex: openShiftOAuth: true notifications: enabled: false -
Apply the Argo CD file:
-
Wait for the ArgoCD Redis deployment to be ready:
-
Enable global IP forwarding on the OVN-Kubernetes configuration:
This command enables IP packet forwarding between different networks managed by OVN-Kubernetes.
Verification
-
Verify that the
MultiClusterEngineinstance is created: -
Verify that the hosted control plane component is enabled:
$ oc get multiclusterengine mce -o jsonpath='{.spec.overrides.components[?(@.name=="hypershift")].enabled}{"\n"}'Note
The manual installation procedure creates
mcewithhypershiftenabled. If the previous command returnstrue, no further action is needed. If it returnsfalseor an empty result, you must enable thehypershiftcomponent before DPU provisioning can succeed.To enable the
hypershiftcomponent on theMultiClusterEngineresource, use the following steps. In current multicluster engine Operator versions the component is namedhypershift. Earlier versions usehypershift-preview.-
If the result is empty, check whether the
componentslist exists:If the list exists but has no
hypershiftentry, append it without replacing the other components:$ oc patch multiclusterengine mce --type=json \ -p='[{"op":"add","path":"/spec/overrides/components/-","value":{"name":"hypershift","enabled":true}}]'If the list is absent, edit the resource instead and create
spec.overrides.componentswith an entry namedhypershiftset toenabled: true. -
If the result is
false, an entry exists but is disabled. Edit the resource and set thehypershiftcomponent toenabled: true:Warning
Do not use
oc patch --type=mergeto enable the component, because a merge patch replaces the entirecomponentsarray and removes the other components. Use the JSONaddpatch only when thecomponentslist exists. Otherwise, useoc edit.
-
-
Verify that the
NodeFeatureDiscoveryinstance andNodeFeatureRuleare configured: -
Verify that the
MetalLBinstance was created: -
Verify that the Argo CD pods are running:
-
Verify that IP forwarding is set to
Global: