Validate traffic and configure telemetry
After provisioning the DPUs and verifying system readiness, validate end-to-end traffic flow and configure DPU telemetry observability.
Deploy traffic test pods and services
You can deploy traffic test pods and services across the management cluster to validate end-to-end connectivity through the DPU data plane. The test workloads include a server pod on a control plane node and worker pods on DPU-enabled nodes, with both standard and host-network configurations.
These test workloads use the nicolaka/netshoot container image, a community networking-troubleshooting image that is not officially supported by Red Hat or NVIDIA. Use it only for connectivity validation and testing, not in production workloads.
Prerequisites
- You have installed the DPF Operator and provisioned the DPU hosted cluster.
- At least one DPU-enabled worker node is available.
- You have access to the management cluster as a user with the
cluster-adminrole.
Procedure
-
Create a file named
traffic-pods.yamlwith the following content:# 1. Namespace---apiVersion: v1kind: Namespacemetadata:name: workload# 2. SCC RoleBinding (grants 'default' ServiceAccount in 'workload' NS access to 'privileged' SCC)---apiVersion: rbac.authorization.k8s.io/v1kind: RoleBindingmetadata:name: privileged-scc-default-sanamespace: workloadsubjects:- kind: ServiceAccountname: defaultnamespace: workloadroleRef:kind: ClusterRolename: system:openshift:scc:privilegedapiGroup: rbac.authorization.k8s.io# 3. Deployments and Services# Deployment: traffic-test-master---apiVersion: apps/v1kind: Deploymentmetadata:name: traffic-test-masternamespace: workloadlabels:app: traffic-test-masterspec:replicas: 1selector:matchLabels:app: traffic-test-mastertemplate:metadata:labels:app: traffic-test-masterspec:topologySpreadConstraints:- maxSkew: 1topologyKey: kubernetes.io/hostnamewhenUnsatisfiable: DoNotSchedulelabelSelector:matchLabels:app: traffic-test-masternodeSelector:node-role.kubernetes.io/control-plane: ""tolerations:- key: node-role.kubernetes.io/masteroperator: Existseffect: NoSchedule- key: node-role.kubernetes.io/control-planeoperator: Existseffect: NoSchedulecontainers:- name: nginxsecurityContext:privileged: truecapabilities:add:- NET_ADMINimage: nicolaka/netshootcommand: ["nc", "-kl", "5000"]ports:- containerPort: 5000name: tcp-serverresources:requests:cpu: 1memory: 1Gilimits:cpu: 1memory: 1Gi---# Service: traffic-test-masterapiVersion: v1kind: Servicemetadata:name: traffic-test-masternamespace: workloadlabels:app: traffic-test-masterspec:selector:app: traffic-test-masterports:- protocol: TCPport: 5000targetPort: 5000---# Service: traffic-test-master-nodeportapiVersion: v1kind: Servicemetadata:name: traffic-test-master-nodeportnamespace: workloadlabels:app: traffic-test-masterspec:type: NodePortselector:app: traffic-test-masterports:- protocol: TCPport: 5000targetPort: 5000---# Deployment: traffic-test-workerapiVersion: apps/v1kind: Deploymentmetadata:name: traffic-test-workernamespace: workloadlabels:app: traffic-test-workerspec:replicas: 1selector:matchLabels:app: traffic-test-workertemplate:metadata:labels:app: traffic-test-workerspec:topologySpreadConstraints:- maxSkew: 1topologyKey: kubernetes.io/hostnamewhenUnsatisfiable: DoNotSchedulelabelSelector:matchLabels:app: traffic-test-workernodeSelector:feature.node.kubernetes.io/dpu-enabled: ""containers:- name: nginxsecurityContext:privileged: truecapabilities:add:- NET_ADMINimage: nicolaka/netshootcommand: ["nc", "-kl", "5000"]ports:- containerPort: 5000name: tcp-serverresources:requests:cpu: 16memory: 6Gilimits:cpu: 16memory: 6Gi---# Service: traffic-test-workerapiVersion: v1kind: Servicemetadata:name: traffic-test-workernamespace: workloadlabels:app: traffic-test-workerspec:selector:app: traffic-test-workerports:- protocol: TCPport: 5000targetPort: 5000---# Service: traffic-test-worker-nodeportapiVersion: v1kind: Servicemetadata:name: traffic-test-worker-nodeportnamespace: workloadlabels:app: traffic-test-workerspec:type: NodePortselector:app: traffic-test-workerports:- protocol: TCPport: 5000targetPort: 5000---# Deployment: traffic-test-worker-hostnetworkapiVersion: apps/v1kind: Deploymentmetadata:name: traffic-test-worker-hostnetworknamespace: workloadlabels:app: traffic-test-worker-hostnetworkspec:replicas: 1selector:matchLabels:app: traffic-test-worker-hostnetworktemplate:metadata:labels:app: traffic-test-worker-hostnetworkspec:topologySpreadConstraints:- maxSkew: 1topologyKey: kubernetes.io/hostnamewhenUnsatisfiable: DoNotSchedulelabelSelector:matchLabels:app: traffic-test-worker-hostnetworknodeSelector:feature.node.kubernetes.io/dpu-enabled: ""hostNetwork: truecontainers:- name: nginxsecurityContext:privileged: truecapabilities:add:- NET_ADMINimage: nicolaka/netshootcommand: ["nc", "-kl", "5000"]ports:- containerPort: 5000name: tcp-serverresources:requests:cpu: 1memory: 1Gilimits:cpu: 1memory: 1Gi---# Service: traffic-test-worker-hostnetworkapiVersion: v1kind: Servicemetadata:name: traffic-test-worker-hostnetworknamespace: workloadlabels:app: traffic-test-worker-hostnetworkspec:selector:app: traffic-test-worker-hostnetworkports:- protocol: TCPport: 5000targetPort: 5000---# Service: traffic-test-worker-hostnetwork-nodeportapiVersion: v1kind: Servicemetadata:name: traffic-test-worker-hostnetwork-nodeportnamespace: workloadlabels:app: traffic-test-worker-hostnetworkspec:type: NodePortselector:app: traffic-test-worker-hostnetworkports:- protocol: TCPport: 5000targetPort: 5000The manifest creates the following resources:
-
A
workloadnamespace for the test pods. -
A
RoleBindingresource that grants thedefaultservice account in theworkloadnamespace access to theprivilegedsecurity context constraint. -
A
traffic-test-masterdeployment andClusterIPandNodePortservices on a control plane node. -
A
traffic-test-workerdeployment andClusterIPandNodePortservices on DPU-enabled worker nodes. -
A
traffic-test-worker-hostnetworkdeployment that uses host networking on DPU-enabled worker nodes, withClusterIPandNodePortservices.noteThe worker deployments set
replicas: 1for a single DPU worker node. Set the replica count of thetraffic-test-workerandtraffic-test-worker-hostnetworkdeployments to the number of DPU-enabled worker nodes so that thetopologySpreadConstraintsplace one pod on each node.
-
-
Apply the manifest:
$ oc apply -f traffic-pods.yaml
Verification
-
Verify that the test pods are running:
$ oc get pods -n workload -o wideExample outputNAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATEStraffic-test-master-7448bb5cc-mdftd 1/1 Running 0 2m10s 10.129.0.145 master-2 <none> <none>traffic-test-worker-776486fb68-krz54 1/1 Running 0 2m10s 10.128.2.9 host-worker1 <none> <none>traffic-test-worker-776486fb68-lz8vf 1/1 Running 0 2m10s 10.131.0.9 host-worker2 <none> <none>traffic-test-worker-hostnetwork-596d569d99-cjpns 1/1 Running 0 2m10s 10.0.110.11 host-worker1 <none> <none>traffic-test-worker-hostnetwork-596d569d99-x6m7r 1/1 Running 0 2m10s 10.0.110.12 host-worker2 <none> <none>Confirm that the
traffic-test-masterpod is on a control plane node and that thetraffic-test-workerpods are distributed across different DPU-enabled worker nodes. -
Verify that the services are created:
$ oc get svc -n workloadExample outputNAME TYPE CLUSTER-IP EXTERNAL-IP PORT(S) AGEtraffic-test-master ClusterIP 172.30.102.123 <none> 5000/TCP 13mtraffic-test-master-nodeport NodePort 172.30.98.22 <none> 5000:31368/TCP 13mtraffic-test-worker ClusterIP 172.30.187.147 <none> 5000/TCP 13mtraffic-test-worker-hostnetwork ClusterIP 172.30.122.242 <none> 5000/TCP 13mtraffic-test-worker-hostnetwork-nodeport NodePort 172.30.122.214 <none> 5000:30209/TCP 13mtraffic-test-worker-nodeport NodePort 172.30.108.72 <none> 5000:32570/TCP 13m
Run traffic validation tests
You can run connectivity tests between the traffic test pods and services to verify that the DPU services and service chains are configured correctly. A successful test confirms that end-to-end traffic flows through the DPU data plane as expected.
Prerequisites
- The traffic test pods and services are deployed in the
workloadnamespace and all pods are in aRunningstate. - You have access to the management cluster as a user with the
cluster-adminrole.
Procedure
-
Run a ping connectivity test between pods on different worker nodes. In the following example, replace
<worker_pod_name>with the name of atraffic-test-workerpod and replace<target_pod_ip>with the IP address of atraffic-test-workerpod on a different worker node:$ oc -n workload exec -it <worker_pod_name> -- ping -c 4 <target_pod_ip>Example outputPING 10.131.0.9 (10.131.0.9) 56(84) bytes of data.64 bytes from 10.131.0.9: icmp_seq=1 ttl=62 time=1.61 ms64 bytes from 10.131.0.9: icmp_seq=2 ttl=62 time=0.876 ms64 bytes from 10.131.0.9: icmp_seq=3 ttl=62 time=0.510 ms64 bytes from 10.131.0.9: icmp_seq=4 ttl=62 time=0.421 ms--- 10.131.0.9 ping statistics ---4 packets transmitted, 4 received, 0% packet loss, time 3028msrtt min/avg/max/mdev = 0.421/0.853/1.606/0.466 msVerify that all 4 packets are received with 0% packet loss.
-
Run a service connectivity test from a worker pod to a service on a control plane node. In the following example, replace
<worker_pod_name>with the name of atraffic-test-workerpod and replace<service_cluster_ip>with the cluster IP address of thetraffic-test-masterservice:$ oc -n workload exec -it <worker_pod_name> -- nc -vz <service_cluster_ip> 5000A
succeededmessage confirms that the service is reachable through the DPU-accelerated network. -
Run a service connectivity test from a worker pod to another worker pod. In the following example, replace
<worker_pod_name>with the name of atraffic-test-workerpod and replace<service_cluster_ip>with the cluster IP address of thetraffic-test-workerservice:$ oc -n workload exec -it <worker_pod_name> -- nc -vz <service_cluster_ip> 5000A
succeededmessage confirms end-to-end connectivity through the DPU-accelerated service chain between worker pods. -
Optional: Run an external service connectivity test. In the following example, replace
<worker_pod_name>with the name of atraffic-test-workerpod, replace<node_ip>with the IP address of a cluster node, and replace<nodeport>with the NodePort for one of the services:$ oc -n workload exec -it <worker_pod_name> -- nc -vz <node_ip> <nodeport>A
succeededmessage confirms that NodePort services are reachable through the DPU networking stack.
DPU telemetry observability with DOCA Telemetry Service
The DOCA Telemetry Service (DTS) exposes DPU hardware telemetry, such as PCIe link speed, uplink throughput, packets, errors, and NIC channel activity, as Prometheus metrics. You can view these metrics by using the OpenShift Container Platform web console or a Grafana dashboard.
In a standard DPF installation, the DTS deployment objects are applied automatically during the postinstallation step. Apply them manually only when you are adding DTS to an existing cluster.
Neither the DPF Operator nor the DTS DPUService installs Grafana on OpenShift Container Platform. Red Hat does not offer a certified Grafana Operator. The community Grafana Operator from OperatorHub is the standard way to run Grafana on OpenShift Container Platform.
DTS runs on every DPU in the hosted cluster and collects counters from sysfs and ethtool providers. OpenShift Container Platform includes a built-in Prometheus instance, so you do not need to deploy a separate monitoring stack to scrape DTS metrics.
How DPF exposes DTS metrics to the management cluster
DTS runs on the DPU hosted cluster, but Prometheus runs on the management cluster. DPF bridges this gap with a built-in port-mirroring mechanism.
When a DPUService resource declares a port in its configPorts field, DPF performs the following actions:
- Publishes the service port as a
NodePorton the DPU hosted cluster. - Creates a mirror
Serviceon the management cluster, labeled withdpu.nvidia.com/exposed-port-for-dpucluster.
The management-cluster Prometheus then scrapes the mirror service. This mechanism requires no additional configuration beyond the standard DTS deployment objects.
DTS deployment objects
DTS is deployed through three standard DPF resources:
DPUServiceTemplate- Defines the Helm chart for the DOCA Telemetry Service, the DTS container image, and the metrics port. The
configMapData.prometheus.portfield is set to9189. DPUServiceConfiguration- Declares the service port
httpserverport: 9189underconfigPorts. This declaration triggers the management-cluster port-mirroring mechanism described previously. DPUDeployment- References the template and configuration so that DTS is rolled out to the DPUs as a
DaemonSeton the DPU hosted cluster. DTS defaults to thesysfsandethtoolproviders.
Enable user workload monitoring for DTS
OpenShift Container Platform includes Prometheus, but by default it only monitors OpenShift Container Platform platform components. You must enable user workload monitoring so that Prometheus can scrape user namespaces where DPF and DTS run, such as dpf-operator-system.
Prerequisites
- A DPF cluster is deployed with at least one provisioned DPU.
- You have access to the management cluster as a user with the
cluster-adminrole.
Procedure
-
Create a file named
cluster-monitoring-config.yamlwith the following content:apiVersion: v1kind: ConfigMapmetadata:name: cluster-monitoring-confignamespace: openshift-monitoringdata:config.yaml: |enableUserWorkload: truenoteIf the
cluster-monitoring-configConfigMapalready exists with other settings, edit it instead of replacing it, and add only theenableUserWorkload: trueline to the existingconfig.yamldata:$ oc -n openshift-monitoring edit configmap cluster-monitoring-config -
Apply the
ConfigMap:$ oc apply -f cluster-monitoring-config.yaml
Verification
-
Verify that the user workload monitoring pods are running in the
openshift-user-workload-monitoringnamespace:$ oc -n openshift-user-workload-monitoring get podsExample outputNAME READY STATUS RESTARTS AGEprometheus-operator-... 1/1 Running 0 ...prometheus-user-workload-0 ... Running 0 ...thanos-ruler-user-workload-0 ... Running 0 ...Confirm that pods named
prometheus-user-workload,thanos-ruler-user-workload, andprometheus-operatorare all in aRunningstate.
Configure the DTS ServiceMonitor
Create a ServiceMonitor resource to instruct the user workload monitoring Prometheus instance to scrape the DOCA Telemetry Service (DTS) metrics endpoint. The ServiceMonitor selects the mirrored DTS service in the dpf-operator-system namespace and scrapes its /metrics path on the httpserverport every 30 seconds.
Prerequisites
- User workload monitoring is enabled in OpenShift Container Platform.
- The DTS
DPUServiceConfigurationandDPUDeploymentresources are applied. For details, see "DPU telemetry observability with DTS". - You have access to the management cluster as a user with the
cluster-adminrole.
Procedure
- Create a file named
dts-servicemonitor.yamlwith the following content:apiVersion: monitoring.coreos.com/v1kind: ServiceMonitormetadata:name: doca-telemetry-service-monitornamespace: dpf-operator-systemspec:selector:matchExpressions:- key: dpu.nvidia.com/dpuservice-nameoperator: Existsendpoints:- port: httpserverportinterval: 30spath: /metricsrelabelings:- sourceLabels:- __meta_kubernetes_service_label_dpu_nvidia_com_dpuservice_nameregex: doca-telemetry-service.*action: keepnamespaceSelector:matchNames:- dpf-operator-system - Apply the
ServiceMonitor:$ oc apply -f dts-servicemonitor.yaml
Verification
-
Verify that the
ServiceMonitoris created in thedpf-operator-systemnamespace:$ oc -n dpf-operator-system get servicemonitor doca-telemetry-service-monitorExample outputNAME AGEdoca-telemetry-service-monitor ...
Install the DTS console dashboard
You can install a DTS dashboard that integrates directly into the OpenShift Container Platform web console. This dashboard provides visibility of DPU telemetry metrics without requiring Grafana or additional tools.
Prerequisites
- User workload monitoring is enabled in OpenShift Container Platform.
- The DTS
ServiceMonitoris configured and collecting metrics. - You have access to the management cluster as a user with the
cluster-adminrole.
Procedure
- Create a file named
dts-console-dashboard.yamlwith the following content to define a console dashboardConfigMapin theopenshift-config-managednamespace:apiVersion: v1kind: ConfigMapmetadata:name: dpf-dts-console-dashboardnamespace: openshift-config-managedlabels:console.openshift.io/dashboard: "true"data:doca-dpu-telemetry-dts.json: |{"title": "DOCA DPU Telemetry (DTS)","uid": "doca-dpu-telemetry-dts-console","editable": false,"schemaVersion": 16,"tags": ["dpf", "dpu", "dts", "telemetry"],"timezone": "browser","time": {"from": "now-1h", "to": "now"},"refresh": "30s","templating": {"list": []},"rows": [{"title": "PCIe / Link","showTitle": true,"height": "250px","panels": [{"type": "graph", "title": "PCIe Link Speed (GT/s)", "span": 6,"datasource": "prometheus", "nullPointMode": "null","legend": {"show": true},"yaxes": [{"format": "none", "show": true}, {"format": "none", "show": false}],"targets": [{"refId": "A", "format": "time_series", "intervalFactor": 2,"expr": "current_link_speed", "legendFormat": "{{source}} {{hca}}"}]},{"type": "graph", "title": "PCIe Link Width (lanes)", "span": 6,"datasource": "prometheus", "nullPointMode": "null","legend": {"show": true},"yaxes": [{"format": "none", "show": true}, {"format": "none", "show": false}],"targets": [{"refId": "A", "format": "time_series", "intervalFactor": 2,"expr": "current_link_width", "legendFormat": "{{source}} {{hca}}"}]}]},{"title": "Uplink Throughput (p0/p1)","showTitle": true,"height": "250px","panels": [{"type": "graph", "title": "Uplink RX (bits/s)", "span": 6,"datasource": "prometheus", "nullPointMode": "null","legend": {"show": true, "values": true, "avg": true, "max": true, "alignAsTable": true, "rightSide": true},"yaxes": [{"format": "bps", "show": true}, {"format": "bps", "show": false}],"targets": [{"refId": "A", "format": "time_series", "intervalFactor": 2,"expr": "sum by(source)(rate({__name__=~\"p[01]_eth_rx_bytes\"}[5m])) * 8","legendFormat": "{{source}}"}]},{"type": "graph", "title": "Uplink TX (bits/s)", "span": 6,"datasource": "prometheus", "nullPointMode": "null","legend": {"show": true, "values": true, "avg": true, "max": true, "alignAsTable": true, "rightSide": true},"yaxes": [{"format": "bps", "show": true}, {"format": "bps", "show": false}],"targets": [{"refId": "A", "format": "time_series", "intervalFactor": 2,"expr": "sum by(source)(rate({__name__=~\"p[01]_eth_tx_bytes\"}[5m])) * 8","legendFormat": "{{source}}"}]}]},{"title": "Uplink Packets & Errors","showTitle": true,"height": "250px","panels": [{"type": "graph", "title": "Uplink packets/s (rx + tx)", "span": 6,"datasource": "prometheus", "nullPointMode": "null","legend": {"show": true, "alignAsTable": true, "rightSide": true},"yaxes": [{"format": "pps", "show": true}, {"format": "pps", "show": false}],"targets": [{"refId": "A", "format": "time_series", "intervalFactor": 2,"expr": "sum by(source)(rate({__name__=~\"p[01]_eth_rx_packets\"}[5m]))","legendFormat": "{{source}} rx"},{"refId": "B", "format": "time_series", "intervalFactor": 2,"expr": "sum by(source)(rate({__name__=~\"p[01]_eth_tx_packets\"}[5m]))","legendFormat": "{{source}} tx"}]},{"type": "graph", "title": "Uplink errors & drops/s", "span": 6,"datasource": "prometheus", "nullPointMode": "null","legend": {"show": true, "alignAsTable": true, "rightSide": true},"yaxes": [{"format": "cps", "show": true}, {"format": "cps", "show": false}],"targets": [{"refId": "A", "format": "time_series", "intervalFactor": 2,"expr": "sum by(source)(rate({__name__=~\"p[01]_eth_rx_errors\"}[5m])) + sum by(source)(rate({__name__=~\"p[01]_eth_tx_errors\"}[5m])) + sum by(source)(rate({__name__=~\"p[01]_eth_rx_dropped\"}[5m])) + sum by(source)(rate({__name__=~\"p[01]_eth_tx_dropped\"}[5m])) + sum by(source)(rate({__name__=~\"p[01]_eth_rx_crc_errors\"}[5m]))","legendFormat": "{{source}}"}]}]},{"title": "NIC Channel Activity","showTitle": true,"height": "250px","panels": [{"type": "graph", "title": "NIC channel poll/s", "span": 6,"datasource": "prometheus", "nullPointMode": "null","legend": {"show": true, "alignAsTable": true, "rightSide": true},"yaxes": [{"format": "cps", "show": true}, {"format": "cps", "show": false}],"targets": [{"refId": "A", "format": "time_series", "intervalFactor": 2,"expr": "sum by(source)(rate(ch_poll[5m]))", "legendFormat": "{{source}}"}]},{"type": "graph", "title": "NIC channel events/s", "span": 6,"datasource": "prometheus", "nullPointMode": "null","legend": {"show": true, "alignAsTable": true, "rightSide": true},"yaxes": [{"format": "cps", "show": true}, {"format": "cps", "show": false}],"targets": [{"refId": "A", "format": "time_series", "intervalFactor": 2,"expr": "sum by(source)(rate(ch_events[5m]))", "legendFormat": "{{source}}"}]}]}]} - Apply the console dashboard:
$ oc apply -f dts-console-dashboard.yaml
Verification
-
Verify that the dashboard
ConfigMapis created:$ oc -n openshift-config-managed get configmap dpf-dts-console-dashboardExample outputNAME DATA AGEdpf-dts-console-dashboard 1 ... -
Access the dashboard in the OpenShift Container Platform web console:
-
Navigate to Observe → Dashboards.
-
In the Dashboard dropdown menu, select DOCA DPU Telemetry (DTS). The dashboard displays PCIe link speed and width, uplink throughput, packets per second, errors and drops per second, and NIC channel activity, with each DPU as its own line.
noteThe DTS console dashboard renders against the platform Thanos or user workload monitoring Prometheus instance. No Grafana dependency is required for basic DPU telemetry viewing.
-
Install Grafana for DTS metrics visualization
You can install the Grafana Operator and a Grafana instance to provide enhanced visualization for DPU telemetry metrics, including per-DPU filtering and customizable dashboards. You can also deploy a dashboard ConfigMap that adds DTS metrics to the OpenShift Container Platform web console.
Prerequisites
- You have enabled user workload monitoring in OpenShift Container Platform.
- You have configured the DTS
ServiceMonitorand it is collecting metrics. - You have access to the management cluster as a user with the
cluster-adminrole. - You have installed the
ocCLI. - You have installed the
helmCLI.
Procedure
-
Install the Grafana Operator by using Helm:
$ helm upgrade -i grafana-operator oci://ghcr.io/grafana/helm-charts/grafana-operator \--version 5.24.0 \--namespace grafana-operator \--create-namespaceExample outputPulled: ghcr.io/grafana/helm-charts/grafana-operator:5.24.0Digest: sha256:4f69cdaecfed2cc61d4e5f4a8e7142795e9b00997e4bcbd37a8c154a225a2f1fRelease "grafana-operator" has been upgraded. Happy Helming!NAME: grafana-operatorLAST DEPLOYED: Tue Aug 4 08:51:36 2026NAMESPACE: grafana-operatorSTATUS: deployedREVISION: 2TEST SUITE: None -
Grant OpenShift Route permissions to the Grafana Operator: The community Grafana Operator requires additional RBAC permissions to manage OpenShift Container Platform routes. Create a file named
grafana-operator-route-rbac.yamlwith the following content:apiVersion: rbac.authorization.k8s.io/v1kind: ClusterRolemetadata:name: grafana-operator-route-managerrules:- apiGroups:- route.openshift.ioresources:- routes- routes/custom-hostverbs:- create- delete- get- list- patch- update- watch---apiVersion: rbac.authorization.k8s.io/v1kind: ClusterRoleBindingmetadata:name: grafana-operator-route-managerroleRef:apiGroup: rbac.authorization.k8s.iokind: ClusterRolename: grafana-operator-route-managersubjects:- kind: ServiceAccountname: grafana-operatornamespace: grafana-operatorApply the file:
$ oc apply -f grafana-operator-route-rbac.yamlExample outputclusterrole.rbac.authorization.k8s.io/grafana-operator-route-manager createdclusterrolebinding.rbac.authorization.k8s.io/grafana-operator-route-manager created -
Create Grafana RBAC for Prometheus access: Create a
ServiceAccountwith a long-lived token and bind it to thecluster-monitoring-viewClusterRoleso Grafana can query the platform Prometheus:apiVersion: v1kind: ServiceAccountmetadata:name: grafana-prometheus-readernamespace: dpf-operator-system---apiVersion: v1kind: Secretmetadata:name: grafana-prometheus-reader-tokennamespace: dpf-operator-systemannotations:kubernetes.io/service-account.name: grafana-prometheus-readertype: kubernetes.io/service-account-token---apiVersion: rbac.authorization.k8s.io/v1kind: ClusterRoleBindingmetadata:name: grafana-prometheus-reader-cluster-monitoringroleRef:apiGroup: rbac.authorization.k8s.iokind: ClusterRolename: cluster-monitoring-viewsubjects:- kind: ServiceAccountname: grafana-prometheus-readernamespace: dpf-operator-systemApply the YAML:
$ oc apply -f grafana-rbac.yaml -
Deploy the Grafana instance with control-plane scheduling:
apiVersion: grafana.integreatly.org/v1beta1kind: Grafanametadata:name: dpf-grafananamespace: dpf-operator-systemlabels:dashboards: "dpf-grafana"spec:route:spec:port:targetPort: grafanatls:termination: edgeconfig:log:mode: "console"level: "info"auth.anonymous:enabled: "true"org_role: "Viewer"security:admin_user: "admin"admin_password: "admin"deployment:spec:template:spec:nodeSelector:node-role.kubernetes.io/control-plane: ""tolerations:- key: node-role.kubernetes.io/masteroperator: Existseffect: NoSchedule- key: node-role.kubernetes.io/control-planeoperator: Existseffect: NoScheduleApply the YAML:
$ oc apply -f grafana-cr.yamlwarningThe default credentials (
admin/admin) are suitable for lab environments only. Change theadmin_passwordfor non-lab deployments. -
Configure the Prometheus datasource:
apiVersion: grafana.integreatly.org/v1beta1kind: GrafanaDatasourcemetadata:name: prometheusnamespace: dpf-operator-systemspec:instanceSelector:matchLabels:dashboards: "dpf-grafana"valuesFrom:- targetPath: "secureJsonData.httpHeaderValue1"valueFrom:secretKeyRef:name: grafana-prometheus-reader-tokenkey: tokendatasource:name: prometheustype: prometheusuid: prometheusaccess: proxyurl: https://thanos-querier.openshift-monitoring.svc.cluster.local:9091isDefault: truejsonData:tlsSkipVerify: truehttpHeaderName1: "Authorization"timeInterval: "30s"secureJsonData:httpHeaderValue1: "Bearer ${token}"Apply the YAML:
$ oc apply -f grafana-datasource.yaml -
Create a file named
dts-grafana-dashboard.yamlwith the following content:apiVersion: v1kind: ConfigMapmetadata:name: dpf-dts-grafana-dashboardnamespace: dpf-operator-systemlabels:app.kubernetes.io/part-of: dpfdata:doca-dpu-telemetry-dts.json: |{"title": "DOCA DPU Telemetry (DTS)","uid": "doca-dpu-telemetry-dts","tags": ["dpf", "dpu", "dts", "telemetry"],"timezone": "browser","schemaVersion": 39,"editable": true,"time": {"from": "now-1h", "to": "now"},"refresh": "30s","templating": {"list": [{"name": "source","label": "DPU (source)","type": "query","datasource": {"type": "prometheus", "uid": "prometheus"},"query": "label_values(current_link_speed, source)","refresh": 2,"includeAll": true,"multi": true,"current": {"text": "All", "value": "$__all"},"sort": 1}]},"panels": [{"type": "stat","title": "PCIe Link Speed (GT/s)","datasource": {"type": "prometheus", "uid": "prometheus"},"gridPos": {"h": 4, "w": 6, "x": 0, "y": 0},"fieldConfig": {"defaults": {"unit": "none"}, "overrides": []},"options": {"reduceOptions": {"calcs": ["lastNotNull"]}, "colorMode": "value", "graphMode": "none"},"targets": [{"refId": "A", "datasource": {"type": "prometheus", "uid": "prometheus"},"expr": "current_link_speed{source=~\"$source\"}", "legendFormat": "{{source}} {{hca}}"}]},{"type": "stat","title": "PCIe Link Width (lanes)","datasource": {"type": "prometheus", "uid": "prometheus"},"gridPos": {"h": 4, "w": 6, "x": 6, "y": 0},"fieldConfig": {"defaults": {"unit": "none"}, "overrides": []},"options": {"reduceOptions": {"calcs": ["lastNotNull"]}, "colorMode": "value", "graphMode": "none"},"targets": [{"refId": "A", "datasource": {"type": "prometheus", "uid": "prometheus"},"expr": "current_link_width{source=~\"$source\"}", "legendFormat": "{{source}} {{hca}}"}]},{"type": "stat","title": "Max PCIe Link Speed (GT/s)","datasource": {"type": "prometheus", "uid": "prometheus"},"gridPos": {"h": 4, "w": 6, "x": 12, "y": 0},"fieldConfig": {"defaults": {"unit": "none"}, "overrides": []},"options": {"reduceOptions": {"calcs": ["lastNotNull"]}, "colorMode": "value", "graphMode": "none"},"targets": [{"refId": "A", "datasource": {"type": "prometheus", "uid": "prometheus"},"expr": "max_link_speed{source=~\"$source\"}", "legendFormat": "{{source}} {{hca}}"}]},{"type": "stat","title": "Max PCIe Link Width (lanes)","datasource": {"type": "prometheus", "uid": "prometheus"},"gridPos": {"h": 4, "w": 6, "x": 18, "y": 0},"fieldConfig": {"defaults": {"unit": "none"}, "overrides": []},"options": {"reduceOptions": {"calcs": ["lastNotNull"]}, "colorMode": "value", "graphMode": "none"},"targets": [{"refId": "A", "datasource": {"type": "prometheus", "uid": "prometheus"},"expr": "max_link_width{source=~\"$source\"}", "legendFormat": "{{source}} {{hca}}"}]},{"type": "timeseries","title": "Uplink RX throughput (p0/p1)","datasource": {"type": "prometheus", "uid": "prometheus"},"gridPos": {"h": 8, "w": 12, "x": 0, "y": 4},"fieldConfig": {"defaults": {"unit": "bps", "custom": {"drawStyle": "line", "fillOpacity": 10}}, "overrides": []},"options": {"legend": {"displayMode": "table", "placement": "bottom", "calcs": ["mean", "max"]}, "tooltip": {"mode": "multi"}},"targets": [{"refId": "A", "datasource": {"type": "prometheus", "uid": "prometheus"},"expr": "sum by(source)(rate({__name__=~\"p[01]_eth_rx_bytes\", source=~\"$source\"}[5m])) * 8","legendFormat": "{{source}}"}]},{"type": "timeseries","title": "Uplink TX throughput (p0/p1)","datasource": {"type": "prometheus", "uid": "prometheus"},"gridPos": {"h": 8, "w": 12, "x": 12, "y": 4},"fieldConfig": {"defaults": {"unit": "bps", "custom": {"drawStyle": "line", "fillOpacity": 10}}, "overrides": []},"options": {"legend": {"displayMode": "table", "placement": "bottom", "calcs": ["mean", "max"]}, "tooltip": {"mode": "multi"}},"targets": [{"refId": "A", "datasource": {"type": "prometheus", "uid": "prometheus"},"expr": "sum by(source)(rate({__name__=~\"p[01]_eth_tx_bytes\", source=~\"$source\"}[5m])) * 8","legendFormat": "{{source}}"}]},{"type": "timeseries","title": "Uplink packets/s (p0/p1 rx+tx)","datasource": {"type": "prometheus", "uid": "prometheus"},"gridPos": {"h": 8, "w": 12, "x": 0, "y": 12},"fieldConfig": {"defaults": {"unit": "pps", "custom": {"drawStyle": "line", "fillOpacity": 10}}, "overrides": []},"options": {"legend": {"displayMode": "table", "placement": "bottom", "calcs": ["mean", "max"]}, "tooltip": {"mode": "multi"}},"targets": [{"refId": "A", "datasource": {"type": "prometheus", "uid": "prometheus"},"expr": "sum by(source)(rate({__name__=~\"p[01]_eth_rx_packets\", source=~\"$source\"}[5m]))","legendFormat": "{{source}} rx"},{"refId": "B", "datasource": {"type": "prometheus", "uid": "prometheus"},"expr": "sum by(source)(rate({__name__=~\"p[01]_eth_tx_packets\", source=~\"$source\"}[5m]))","legendFormat": "{{source}} tx"}]},{"type": "timeseries","title": "Uplink errors & drops/s (p0/p1)","datasource": {"type": "prometheus", "uid": "prometheus"},"gridPos": {"h": 8, "w": 12, "x": 12, "y": 12},"fieldConfig": {"defaults": {"unit": "cps", "custom": {"drawStyle": "line", "fillOpacity": 10}}, "overrides": []},"options": {"legend": {"displayMode": "table", "placement": "bottom", "calcs": ["max"]}, "tooltip": {"mode": "multi"}},"targets": [{"refId": "A", "datasource": {"type": "prometheus", "uid": "prometheus"},"expr": "sum by(source)(rate({__name__=~\"p[01]_eth_rx_errors\", source=~\"$source\"}[5m])) + sum by(source)(rate({__name__=~\"p[01]_eth_tx_errors\", source=~\"$source\"}[5m])) + sum by(source)(rate({__name__=~\"p[01]_eth_rx_dropped\", source=~\"$source\"}[5m])) + sum by(source)(rate({__name__=~\"p[01]_eth_tx_dropped\", source=~\"$source\"}[5m])) + sum by(source)(rate({__name__=~\"p[01]_eth_rx_crc_errors\", source=~\"$source\"}[5m]))","legendFormat": "{{source}}"}]},{"type": "timeseries","title": "NIC channel poll/s","datasource": {"type": "prometheus", "uid": "prometheus"},"gridPos": {"h": 8, "w": 12, "x": 0, "y": 20},"fieldConfig": {"defaults": {"unit": "cps", "custom": {"drawStyle": "line", "fillOpacity": 10}}, "overrides": []},"options": {"legend": {"displayMode": "table", "placement": "bottom", "calcs": ["mean", "max"]}, "tooltip": {"mode": "multi"}},"targets": [{"refId": "A", "datasource": {"type": "prometheus", "uid": "prometheus"},"expr": "rate(ch_poll{source=~\"$source\"}[5m])","legendFormat": "{{source}} {{device_name}}"}]},{"type": "timeseries","title": "NIC channel events/s","datasource": {"type": "prometheus", "uid": "prometheus"},"gridPos": {"h": 8, "w": 12, "x": 12, "y": 20},"fieldConfig": {"defaults": {"unit": "cps", "custom": {"drawStyle": "line", "fillOpacity": 10}}, "overrides": []},"options": {"legend": {"displayMode": "table", "placement": "bottom", "calcs": ["mean", "max"]}, "tooltip": {"mode": "multi"}},"targets": [{"refId": "A", "datasource": {"type": "prometheus", "uid": "prometheus"},"expr": "rate(ch_events{source=~\"$source\"}[5m])","legendFormat": "{{source}} {{device_name}}"}]}]}---apiVersion: grafana.integreatly.org/v1beta1kind: GrafanaDashboardmetadata:name: doca-dpu-telemetry-dtsnamespace: dpf-operator-systemspec:instanceSelector:matchLabels:dashboards: "dpf-grafana"configMapRef:name: dpf-dts-grafana-dashboardkey: doca-dpu-telemetry-dts.jsonApply the YAML:
$ oc apply -f dts-grafana-dashboard.yaml
Verification
-
Verify that the Grafana Operator is running:
$ oc get pods -n grafana-operatorExample outputNAME READY STATUS RESTARTS AGEgrafana-operator-66d8c8c7b-xyz12 1/1 Running 0 5m42s -
Verify that the Grafana instance is running:
$ oc get grafana -n dpf-operator-systemExample outputNAME AGEdpf-grafana 3m15s -
Verify that the Grafana pod is running on a control plane node:
$ oc get pods -n dpf-operator-system -o wide | grep grafanaExample outputNAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATESgrafana-deployment-7c8b9d-xyz12 1/1 Running 0 2m38s 10.128.0.45 master-node-1 <none> <none> -
Verify the Grafana route exists:
$ oc get route -n dpf-operator-systemExample outputNAME HOST/PORT PATH SERVICES PORT TERMINATION WILDCARDdpf-grafana-route dpf-grafana-route-dpf-operator-system.apps.cluster.example.com dpf-grafana grafana edge None -
Access the Grafana web interface: Get the Grafana URL:
$ echo "https://$(oc -n dpf-operator-system get route dpf-grafana-route -o jsonpath='{.spec.host}')"Open the returned URL in a web browser. Use anonymous access (read-only) or sign in with the default credentials (
admin/admin) for editing capabilities. -
Navigate to the DTS dashboard:
- In Grafana, go to Dashboards and open DOCA DPU Telemetry (DTS).
- Use the DPU (source) dropdown menu to focus on a specific DPU or select All.
- Adjust the time range by using the time-range control on the dashboard toolbar. The dashboard refreshes every 30 seconds.
- Optional: In the OpenShift Container Platform web console, go to Observe → Dashboards and open DOCA DPU Telemetry (DTS) to view the console-integrated dashboard.
- Review the following metrics:
- PCIe status: current and maximum link speed and width
- Throughput: receive (RX) and transmit (TX) data rates for uplink ports
p0andp1 - Packet rates: packets-per-second statistics with RX and TX breakdown
- Error monitoring: combined error, drop, and CRC error rates
View DTS metrics and dashboards
After you configure the DTS ServiceMonitor, you can view DPU telemetry metrics by using the OpenShift Container Platform web console, PromQL queries, or Grafana dashboards.
Prerequisites
- You have configured the DTS
ServiceMonitor. - You have access to the management cluster as a user with the
cluster-adminrole. - You have installed the
ocCLI. - Optional: You have installed Grafana for DTS metrics visualization.
Procedure
-
Verify that the DTS
DPUServiceis ready on the management cluster. The object name carries a generated suffix, so select it by its stable label:$ oc -n dpf-operator-system get dpuservice \-l svc.dpu.nvidia.com/dpudeployment-service=doca-telemetry-serviceExample outputNAME READY PHASE AGEdoca-telemetry-service-89p28 True Success ...A status of
READY: TrueandPHASE: Successconfirms that DTS is deployed and running. -
View metrics in the OpenShift Container Platform web console.
noteThe OpenShift Container Platform web console reads from the cluster Prometheus instance through Thanos and user workload monitoring. Grafana is not required for basic metric viewing.
To run an ad hoc query, go to Observe → Metrics in the web console, enter a DTS
PromQLquery, and click Run queries.Example query to validate DTS metrics:
current_link_speed{job=~"doca-telemetry-service.*"}You should get one series per DPU. Hover over a line to see its labels. Note the
sourcelabel, which is the DPU node name that identifies each DPU.DTS
PromQLqueriesQuery Description current_link_speed{job=~"doca-telemetry-service.*"}Returns the current PCIe link speed for each DPU. Each series includes a sourcelabel that identifies the DPU node name.rate(p0_eth_rx_bytes{job=~"doca-telemetry-service.*"}[5m]) * 8Calculates the uplink receive throughput in bits per second over a 5-minute window. rate(ch_poll{job=~"doca-telemetry-service.*"}[5m])Calculates the NIC channel polling activity rate over a 5-minute window. To view the console dashboard, go to Observe → Dashboards, then in the Dashboard dropdown menu, select DOCA DPU Telemetry (DTS). The dashboard displays PCIe link speed and width, uplink throughput, packets per second, errors and drops per second, and NIC channel activity, with each DPU as its own line.
-
Optional: View metrics in Grafana. Grafana provides richer dashboards with per-DPU dropdown filters and customizable panels. After you install the Grafana Operator and Grafana instance, retrieve the route URL:
$ echo "https://$(oc -n dpf-operator-system get route dpf-grafana-route -o jsonpath='{.spec.host}')"Open the outputted URL in a browser. Anonymous access provides read-only viewer permissions. To edit dashboards, click Sign in and use
admin/adminas the default credentials set in the Grafana custom resource.warningChange the default Grafana credentials for non-lab clusters.
In Grafana, go to Dashboards and open DOCA DPU Telemetry (DTS). Use the DPU (source) dropdown menu to focus on a specific DPU or select All. Adjust the time range by using the time-range control on the dashboard toolbar. The dashboard refreshes every 30 seconds.
-
Optional: Review DPF framework dashboards in Grafana. The DPF Operator installs framework dashboards that track DPU lifecycle and control-plane health separately from the DTS hardware telemetry dashboard. These dashboards are loaded into Grafana automatically through
GrafanaDashboardresources created fromConfigMaps.DPF framework dashboards
Dashboard Description DOCA Platform DPU Fleet Health Fleet-wide DPU health, provisioning state, and version distribution. DOCA Platform DPU Health Detail Per-DPU status, conditions, and history timelines. DOCA Platform Framework State Inventory and readiness of every DPF resource type. DOCA Platform Framework Performance Time for DPF resources to reach their conditions, including reconcile and provisioning timings. Controller Runtime DPF controller internals: CPU and memory usage, reconcile rates, queues, and errors.