Integrating Dynamic Resource Allocation
You can configure Red Hat build of Kueue to manage quota for workloads that use Dynamic Resource Allocation (DRA) to request GPUs. When DRA quota management is configured, Red Hat build of Kueue counts DRA device requests toward quota in the same way that it counts traditional resources such as CPU and memory.
Kueue integration with Dynamic Resource Allocation (DRA) is a Technology Preview feature only. Technology Preview features are not supported with Red Hat production service level agreements (SLAs) and might not be functionally complete. Red Hat does not recommend using them in production. These features provide early access to upcoming product features, enabling customers to test functionality and provide feedback during the development process.
For more information about the support scope of Red Hat Technology Preview features, see Technology Preview Features Support Scope.
If DRA device quota is not configured, Red Hat build of Kueue does not account for GPU requests when admitting workloads, which can result in teams exceeding their GPU allocation.
DRA quota management overview
Dynamic Resource Allocation (DRA) is a Kubernetes framework that provides structured discovery and allocation of specialized hardware such as GPUs. DRA drivers publish device information through ResourceSlice objects, and administrators group devices into categories using DeviceClass objects.
Without Red Hat build of Kueue DRA integration, GPU requests made through DRA are invisible to quota management. Red Hat build of Kueue cannot account for these requests when admitting workloads, which can result in teams exceeding their GPU allocation.
Red Hat build of Kueue provides two approaches for managing DRA device quota:
ResourceClaimTemplate- The default approach. Workloads explicitly reference a
ResourceClaimTemplateobject that defines device requirements. Administrators configuredeviceClassMappingsin the Kueue CR to map eachDeviceClassobject to a logical resource name for quota tracking. Use this approach when workloads need fine-grained control over device selection, such as targeting a specific GPU model or architecture using CEL selectors. Extended resources- A simplified alternative that allows workloads to use standard Kubernetes
resources.requestssyntax, for example,nvidia.com/gpu: "1", instead of explicitly creating DRA objects. When aDeviceClassobject includes thespec.extendedResourceNamefield, the Kubernetes scheduler automatically generatesResourceClaimobjects. Use this approach when you want the simplest possible user experience and backward compatibility with existing workload YAML.noteIf a
DeviceClassobject with theextendedResourceNamefield also appears in adeviceClassMappingsentry, Red Hat build of Kueue uses the mapped logical name from thedeviceClassMappingsentry for quota instead of the extended resource name, unifying quota accounting across both paths.
For clusters with partitionable devices, such as NVIDIA Multi-Instance GPU (MIG), Red Hat build of Kueue can also charge quota in capacity units, such as GPU memory, rather than device count. Partitionable devices use ResourceClaimTemplates objects with CEL selectors to target specific partition profiles, and require administrators to configure counter-based sources in deviceClassMappings. This capability requires OpenShift Container Platform 4.22 or later.
Configuring the resource claim template path
You can configure Red Hat build of Kueue to manage quota for workloads that explicitly reference ResourceClaimTemplate objects. This requires configuring the deviceClassMappings entry in the Red Hat build of Kueue custom resource (CR) and adding the DRA resource to your ClusterQueue object.
Prerequisites
-
You have installed Red Hat build of Kueue by using the Red Hat Build of Kueue Operator.
-
You have created a
Kueuecustom resource (CR). -
Your cluster is running OpenShift Container Platform 4.21 or later.
-
A DRA driver is installed in the cluster, for example,
nvidia-dra-driver. You can verify that the DRA driver is publishing device information by running the following command:$ oc get resourceslicesIf the command returns one or more
ResourceSliceobjects, the DRA driver is running. -
At least one
DeviceClassobject exists in the cluster. You can verify this by running the following command:$ oc get deviceclass
Procedure
-
Use the following command to add a
deviceClassMappingsentry to the Red Hat build of Kueue configuration that maps eachDeviceClassto a logical resource name for quota:$ oc patch kueue cluster -n openshift-kueue-operator --type=merge -p '{"spec": {"config": {"resources": {"deviceClassMappings": [{"name": "nvidia.com/gpu","deviceClassNames": ["gpu.nvidia.com"]}]}}}}'Replace
"nvidia.com/gpu"with the resource name used inClusterQueuequotas andWorkloadstatus.Replace
"gpu.nvidia.com"with one or moreDeviceClassnames that map to this resource.Multiple device classes can map to the same logical resource name. For example, if you have separate device classes for different GPU models but want a single quota pool, as shown in the following example:
resources:deviceClassMappings:- name: nvidia.com/gpudeviceClassNames:- gpu-a100.nvidia.com- gpu-h100.nvidia.com -
Create a file called
rct-queues.yamlthat contains the following content:Example quota configuration for a ResourceClaimTemplate objectapiVersion: kueue.x-k8s.io/v1beta2kind: ResourceFlavormetadata:name: "default-flavor"---apiVersion: kueue.x-k8s.io/v1beta2kind: ClusterQueuemetadata:name: "cluster-queue"spec:namespaceSelector: {}resourceGroups:- coveredResources: ["cpu", "memory", "nvidia.com/gpu"]flavors:- name: "default-flavor"resources:- name: "cpu"nominalQuota: 40- name: "memory"nominalQuota: 200Gi- name: "nvidia.com/gpu"nominalQuota: 8---apiVersion: kueue.x-k8s.io/v1beta2kind: LocalQueuemetadata:namespace: "default"name: "user-queue"spec:clusterQueue: "cluster-queue" -
Apply the
rct-queues.yamlfile:$ oc apply -f rct-queues.yaml -
Create a
ResourceClaimTemplateobject and a workload to verify the configuration. Create a file calledrct-job.yamlby running the following command:$ oc create -f rct-job.yamlExample ResourceClaimTemplate workloadapiVersion: resource.k8s.io/v1kind: ResourceClaimTemplatemetadata:name: my-gpunamespace: defaultspec:spec:devices:requests:- name: gpuexactly:deviceClassName: gpu.nvidia.com---apiVersion: batch/v1kind: Jobmetadata:generateName: rct-test-job-namespace: defaultlabels:kueue.x-k8s.io/queue-name: user-queuespec:template:spec:restartPolicy: NeverresourceClaims:- name: gpuresourceClaimTemplateName: my-gpucontainers:- name: workerimage: registry.k8s.io/e2e-test-images/agnhost:2.53args: ["pause"]resources:claims:- name: gpurequests:cpu: "1"memory: "200Mi"where:
spec.spec.drivers.requests.exactly.deviceClassName:- References the
DeviceClassobject configured in thedeviceClassMappingsentry. metadata.labels.kueue.x-k8s.io/queue-name:- Identifies the local queue to submit the job to.
spec.template.spec.resourceClaims.resourceClaimTemplateName:- References the
ResourceClaimTemplateobject defined above. The template must exist in the same namespace as the job. spec.template.spec.containers.resources.claims.name:- Attaches the resource claim to this container.
Verification
-
Verify that the workload has been created and admitted:
$ oc -n default get workloads -
Verify that a
ResourceClaimobject was created from the template:$ oc -n default get resourceclaimsIf the workload is not admitted, verify the following:
- Check if the namespace is managed by Red Hat build of Kueue:
$ oc label namespace default kueue.openshift.io/managed=true
- The
deviceClassMappingsin theKueueCR maps theDeviceClassobject to the resource name in thecoveredResourcesparameter. - The
ClusterQueueobject has sufficient quota available. - The
ResourceClaimTemplateobject exists in the same namespace as the job.
- Check if the namespace is managed by Red Hat build of Kueue:
Configuring the extended resources path
You can configure Red Hat build of Kueue to manage quota for workloads that request GPUs by using the standard resources.requests syntax, for example, nvidia.com/gpu: "1".
When a DeviceClass object includes the spec.extendedResourceName field, the Kubernetes scheduler automatically generates ResourceClaim objects. This path does not require deviceClassMappings configuration because Red Hat build of Kueue auto-discovers the mapping by indexing DeviceClass objects.
The Red Hat build of Kueue Operator automatically enables the required Red Hat build of Kueue feature gates when it detects the DRAExtendedResource Kubernetes feature gate on the cluster. No manual Red Hat build of Kueue feature gate configuration is required.
To use the extended resources path, you must enable the DRAExtendedResource Kubernetes feature gate. This feature is expected to be generally available in a future OpenShift Container Platform release.
Prerequisites
-
You have cluster administrator permissions.
-
You have installed Red Hat build of Kueue by using the Red Hat Build of Kueue Operator.
-
You have created a
KueueCR. -
Your cluster is running OpenShift Container Platform 4.21 or later.
-
A DRA driver is installed in the cluster, for example,
nvidia-dra-driver. You can verify that the DRA driver is publishing device information by running the following command:$ oc get resourceslicesIf the command returns one or more
ResourceSliceobjects, the DRA driver is running. -
At least one
DeviceClassobject exists in the cluster. You can verify this by running the following command:$ oc get deviceclass -
You have enabled the
DRAExtendedResourceKubernetes feature gate by adding theCustomNoUpgradefeature set to theFeatureGateCR namedcluster, as shown in the following example:apiVersion: config.openshift.io/v1kind: FeatureGatemetadata:name: clusterspec:featureSet: CustomNoUpgradecustomNoUpgrade:enabled:- DRAExtendedResourcewarningEnabling the
CustomNoUpgradefeature set on your cluster cannot be undone and prevents minor version updates. This feature set is not supported on production clusters.
Procedure
-
Verify that the
DeviceClassobject hasspec.extendedResourceNameset by running the following command:$ oc get deviceclass gpu.nvidia.com -o jsonpath='{.spec.extendedResourceName}'Example outputnvidia.com/gpuIf the command does not return a value, add the
extendedResourceNamefield by running the following command:$ oc patch deviceclass gpu.nvidia.com --type=merge -p '{"spec":{"extendedResourceName":"nvidia.com/gpu"}}' -
Create a
ClusterQueueobject that includes the GPU resource in thecoveredResourcesparameter by creating a file calleder-queues.yaml, as shown in the following example:Example quota configuration for extended resourcesapiVersion: kueue.x-k8s.io/v1beta2kind: ResourceFlavormetadata:name: "default-flavor"---apiVersion: kueue.x-k8s.io/v1beta2kind: ClusterQueuemetadata:name: "cluster-queue"spec:namespaceSelector: {}resourceGroups:- coveredResources: ["cpu", "memory", "nvidia.com/gpu"]flavors:- name: "default-flavor"resources:- name: "cpu"nominalQuota: 40- name: "memory"nominalQuota: 200Gi- name: "nvidia.com/gpu"nominalQuota: 8---apiVersion: v1kind: Namespacemetadata:name: team-alabels:kueue.openshift.io/managed: "true"---apiVersion: kueue.x-k8s.io/v1beta2kind: LocalQueuemetadata:namespace: "team-a"name: "user-queue"spec:clusterQueue: "cluster-queue" -
Apply the quota configuration by running the following command:
$ oc apply -f er-queues.yaml -
Create a workload that uses the standard resource request syntax by creating a file called
er-job.yaml, as shown in the following example:Example workload using extended resourcesapiVersion: batch/v1kind: Jobmetadata:name: er-test-jobnamespace: team-alabels:kueue.x-k8s.io/queue-name: user-queuespec:template:spec:containers:- name: workerimage: registry.k8s.io/e2e-test-images/agnhost:2.53args: ["pause"]resources:requests:cpu: "1"memory: "200Mi"nvidia.com/gpu: "1"limits:nvidia.com/gpu: "1"restartPolicy: Neverwhere:
metadata.labels.kueue.x-k8s.io/queue-name- Identifies the local queue to submit the job to.
spec.template.spec.containers.resources.requests.cpu.nvidia.com/gpu- Requests a GPU by using the standard extended resource syntax. No
ResourceClaimTemplateorresourceClaimssection is needed. TheDeviceClassobject with thespec.extendedResourceNamefield causes the Kubernetes scheduler to generate aResourceClaimobject automatically. spec.template.spec.containers.resources.limits.cpu.nvidia.com/gpu- Replace
"1"with a GPU by using the standard extended resource syntax. NoResourceClaimTemplateorresourceClaimssection is needed. TheDeviceClassobject with thespec.extendedResourceNamefield causes the Kubernetes scheduler to generate aResourceClaimobject automatically.
- Create the workload by running the following command:
$ oc apply -f er-job.yaml
Verification
-
Verify that a workload has been created and admitted by running the following command:
$ oc -n team-a get workloadsExample outputNAME QUEUE RESERVED IN ADMITTED AGEjob-er-test-job-4m2x-d3f4g user-queue cluster-queue True 10s -
Verify that a
ResourceClaimwas automatically created by running the following command:$ oc -n team-a get resourceclaimsExample outputNAME STATE AGEer-test-job-jj7vz-extended-resources-bggzk allocated,reserved 24sThe Kubernetes scheduler creates a
ResourceClaimfor each pod that requests an extended resource backed by aDeviceClass.If the workload is not admitted, verify the following:
- The
DRAExtendedResourceKubernetes feature gate is enabled on the cluster. - The
DeviceClasshasspec.extendedResourceNameset. - The
ClusterQueueincludes the extended resource name incoveredResources. - The
ClusterQueuehas sufficient quota available.
- The
Configuring the partitionable devices
You can configure Red Hat build of Kueue to manage quota for partitionable devices based on actual device capacity rather than device count. Partitionable devices, such as NVIDIA Multi-Instance GPU (MIG) capable GPUs, allow a single GPU to be dynamically subdivided into smaller partitions.
When counter-based quota is configured, Red Hat build of Kueue charges quota in capacity units such as GPU memory rather than counting whole devices. For example, a 1g.5gb MIG partition on an A100-40GB charges 4864Mi of GPU memory quota, while a whole GPU charges 40320Mi.
To use partitionable devices, your cluster must be running OpenShift Container Platform 4.22 or later and must have the CustomNoUpgrade feature set enabled with explicit DRAPartitionableDevices gate enablement.
Prerequisites
-
You have cluster administrator permissions.
-
You have installed Red Hat build of Kueue by using the Red Hat Build of Kueue Operator.
-
You have created a
Kueuecustom resource (CR). -
Your cluster is running OpenShift Container Platform 4.22 or later.
-
A DRA driver that publishes
consumesCountersinResourceSliceobjects is installed, for example,nvidia-dra-driver. You can verify that the DRA driver is publishing device information by running the following command:$ oc get resourceslicesIf the command returns one or more
ResourceSliceobjects, the DRA driver is running. -
At least one
DeviceClassobject exists in the cluster. You can verify this by running the following command:$ oc get deviceclass -
MIG is enabled on the GPU hardware.
-
You have enabled the
DRAPartitionableDevicesKubernetes feature gate by adding theCustomNoUpgradefeature set to theFeatureGateCR namedcluster, as shown in the following example:apiVersion: config.openshift.io/v1kind: FeatureGatemetadata:name: clusterspec:featureSet: CustomNoUpgradecustomNoUpgrade:enabled:- DRAPartitionableDeviceswarningEnabling the
CustomNoUpgradefeature set on your cluster cannot be undone and prevents minor version updates. This feature set is not supported on production clusters. For information about enabling feature gates, see "Enabling features using feature gates".
Procedure
-
Verify that your DRA driver publishes counter data by running the following command:
$ oc get resourceslices -o jsonpath='{range .items[*]}{.spec.driver}{"\t"}{range .spec.devices[*]}{.name}: {.consumesCounters}{"\n"}{end}{end}'Example outputgpu.nvidia.com gpu-0: [{"counterSet":"shared","counters":{"memory":{"value":"40Gi"}}}]If the output does not show
consumesCountersdata, verify that your DRA driver version supports partitionable devices and that MIG is enabled on the GPU hardware. -
Configure counter-based quota by adding a
deviceClassMappingsentry with asourcessection to theconfig.resourcessection of the Red Hat build of Kueue CR, as shown in the following example:apiVersion: kueue.openshift.io/v1kind: Kueuemetadata:name: clusternamespace: openshift-kueue-operatorspec:config:resources:deviceClassMappings:- name: gpu.memorydeviceClassNames:- gpu.nvidia.com- mig.nvidia.comsources:- type: Countercounter:name: memorydriver: gpu.nvidia.comdeviceSelector:type: CELcel:expression: "device.driver == 'gpu.nvidia.com'"# ...where:
spec.config.resources.deviceClassMappings.name- The logical resource name used in
ClusterQueuequotas. When counter-based sources are configured, quota is charged in capacity units rather than device count.
spec.config.resources.deviceClassMappings.deviceClassNames- The
DeviceClassnames that map to this resource. Include both the whole-GPU class (gpu.nvidia.com) and the MIG class (mig.nvidia.com). spec.config.resources.deviceClassMappings.sources- Defines how Red Hat build of Kueue computes the quota charge.
spec.config.resources.deviceClassMappings.sources.counter.name- The counter name must match a counter key published by the DRA driver in
ResourceSlicedevices. spec.config.resources.deviceClassMappings.sources.counter.deviceSelector- Scopes which devices are eligible for counter-based quota accounting.
note
The Red Hat build of Kueue Operator automatically enables the required Red Hat build of Kueue feature gates when it detects the
DRAPartitionableDevicesKubernetes feature gate andsourcesare configured indeviceClassMappings. No manual Red Hat build of Kueue feature gate configuration is required.
-
Create a
ClusterQueueobject with counter-based quota. Set the quota in capacity units rather than device count. Create a file calledpd-queues.yamlwith the following content:Example quota configuration for partitionable devicesapiVersion: kueue.x-k8s.io/v1beta2kind: ResourceFlavormetadata:name: "default-flavor"---apiVersion: kueue.x-k8s.io/v1beta2kind: ClusterQueuemetadata:name: "cluster-queue"spec:namespaceSelector: {}resourceGroups:- coveredResources: ["cpu", "memory", "gpu.memory"]flavors:- name: "default-flavor"resources:- name: "cpu"nominalQuota: 40- name: "memory"nominalQuota: 200Gi- name: "gpu.memory"nominalQuota: 800Gi---apiVersion: v1kind: Namespacemetadata:name: team-alabels:kueue.openshift.io/managed: "true"---apiVersion: kueue.x-k8s.io/v1beta2kind: LocalQueuemetadata:namespace: "team-a"name: "user-queue"spec:clusterQueue: "cluster-queue"where:
spec.resourceGroups.coveredResources- The
gpu.memoryentry must match thenamevalue indeviceClassMappings.
spec.resourceGroups.flavors.resources.name- Specify
"gpu.memory"to set the total GPU memory quota. For example,800Giaccommodates twenty A100-40GB GPUs or equivalent MIG partitions.noteWhen
ClusterQueueobjects share a cohort, ensure all queues use the same unit scale for counter resources. Red Hat build of Kueue does not validate unit consistency acrossClusterQueueobjects.
-
Apply the quota configuration by running the following command:
$ oc apply -f pd-queues.yaml -
Create a workload that requests a MIG partition by creating a file called
pd-job.yaml, as shown in the following example:Example workload requesting a MIG partitionapiVersion: resource.k8s.io/v1kind: ResourceClaimTemplatemetadata:namespace: team-aname: gpu-partitionspec:spec:devices:requests:- name: gpuexactly:deviceClassName: mig.nvidia.comcount: 1selectors:- cel:expression: "device.attributes['gpu.nvidia.com'].profile == '1g.5gb'"---apiVersion: batch/v1kind: Jobmetadata:generateName: pd-test-jobnamespace: team-alabels:kueue.x-k8s.io/queue-name: user-queuespec:template:spec:containers:- name: workerimage: registry.k8s.io/e2e-test-images/agnhost:2.53args: ["pause"]resources:claims:- name: gpurequests:cpu: "1"memory: "200Mi"resourceClaims:- name: gpuresourceClaimTemplateName: gpu-partitionrestartPolicy: Neverwhere:
spec.spec.devices.requests.exactly.deviceClassName- References the MIG
DeviceClass.
spec.spec.devices.requests.exactly.selectors.cel.expression:- Selects a specific MIG partition profile. Available profiles depend on the GPU model, for example,
1g.5gb,2g.10gb,3g.20gb, or7g.40gbfor the A100-40GB. metadata.labels.kueue.x-k8s.io/queue-name:- Identifies the local queue to submit the job to.
spec.template.spec.resourceClaims.resourceClaimTemplateName:- References the
ResourceClaimTemplatedefined above. TheResourceClaimTemplatemust exist in the same namespace as the job.
- Create the workload by running the following command:
$ oc create -f pd-job.yaml
Verification
-
Verify that the workload is admitted and that quota was charged in capacity units by running the following command:
$ oc -n team-a get workloads -o jsonpath='{range .items[*]}{.metadata.name}: {.status.admission.podSetAssignments[0].resourceUsage}{"\n"}{end}'Example outputjob-pd-test-job-xxxxx: {"cpu":"1","gpu.memory":"5100273664","memory":"200Mi"}The
gpu.memoryvalue reflects the actual memory capacity of the requested MIG partition rather than a device count of1. -
If the workload is not admitted, verify the following:
- The
DRAPartitionableDevicesKubernetes feature gate is enabled on the cluster. - The
namevalue of thedeviceClassMappingsobject matches the resource name incoveredResources. - The
counter.nameinsourcesmatches a counter key in theResourceSliceobjects. - The
ClusterQueuehas sufficient GPU memory quota for the requested partition size. - MIG is enabled on the GPU hardware.
- The
Additional resources