Run Workloads With DRA Devices
This page shows you how to run workloads that request hardware devices (such as GPUs) managed by Dynamic Resource Allocation (DRA) in a Kubernetes cluster with Kueue enabled. The examples use a batch Job, but the same approach works with any workload type that Kueue supports.
The intended audience for this page are batch users.
For conceptual details about how Kueue handles DRA resources, see Dynamic Resource Allocation concepts.
Before you begin
Make sure the following conditions are met:
- A Kubernetes cluster is running.
- The kubectl command-line tool has communication with your cluster.
- Kueue is installed.
- The cluster has quotas configured
with DRA resources included in the
ClusterQueue. - Your administrator has set up DRA support in Kueue.
0. Identify the queues available in your namespace
Run the following command to list the LocalQueues available in your namespace.
kubectl -n default get localqueues
The output is similar to the following:
NAME CLUSTERQUEUE PENDING WORKLOADS
user-queue cluster-queue 0
The ClusterQueue defines the quotas for the Queue.
1. Define the workload
Running a workload with DRA devices is similar to
running a regular Job. You must set the
kueue.x-k8s.io/queue-name label to select the LocalQueue you want to
submit the workload to.
There are several ways to request DRA devices, depending on how your administrator has configured the cluster. Choose the approach that matches your setup.
Using a ResourceClaimTemplate
Use this approach when you need to explicitly describe the device you want.
Create a ResourceClaimTemplate and reference it from the workload:
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
namespace: default
name: single-gpu
spec:
spec:
devices:
requests:
- name: gpu
exactly:
deviceClassName: gpu.example.com
---
apiVersion: batch/v1
kind: Job
metadata:
generateName: sample-dra-rct-job-
namespace: default
labels:
kueue.x-k8s.io/queue-name: user-queue
spec:
template:
spec:
containers:
- name: dummy-job
image: registry.k8s.io/e2e-test-images/agnhost:2.53
args: ["pause"]
resources:
claims:
- name: gpu
requests:
cpu: "1"
memory: "200Mi"
resourceClaims:
- name: gpu
resourceClaimTemplateName: single-gpu
restartPolicy: Never
Using a firstAvailable request
Note
This feature requires theKueueDRAIntegrationPrioritizedList feature gate,
which is disabled by default in v0.20.If your administrator has
set up firstAvailable requests,
list the alternative device classes in order of preference. They must use
ExactCount, ask for the same count, and map to one logical resource:
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
namespace: default
name: a100-or-mig
spec:
spec:
devices:
requests:
- name: gpu
firstAvailable:
- name: a100
deviceClassName: a100.example.com
- name: mig
deviceClassName: a100-mig.example.com
---
apiVersion: batch/v1
kind: Job
metadata:
generateName: sample-dra-firstavailable-job-
namespace: default
labels:
kueue.x-k8s.io/queue-name: user-queue
spec:
template:
spec:
containers:
- name: dummy-job
image: registry.k8s.io/e2e-test-images/agnhost:2.53
args: ["pause"]
resources:
claims:
- name: gpu
requests:
cpu: "1"
memory: "200Mi"
resourceClaims:
- name: gpu
resourceClaimTemplateName: a100-or-mig
restartPolicy: Never
Kueue charges this one-Pod Job example.com/gpu: 1, whichever alternative the
kube-scheduler
allocates.
Using extended resources
Use this approach when a DeviceClass with spec.extendedResourceName exists
in the cluster. You request devices using the standard resources.requests
syntax, just like CPU or memory. No ResourceClaimTemplate is needed:
apiVersion: batch/v1
kind: Job
metadata:
generateName: sample-dra-extended-job-
namespace: default
labels:
kueue.x-k8s.io/queue-name: user-queue
spec:
template:
spec:
containers:
- name: dummy-job
image: registry.k8s.io/e2e-test-images/agnhost:2.53
args: ["pause"]
resources:
requests:
cpu: "1"
memory: "200Mi"
example.com/gpu: "1"
limits:
example.com/gpu: "1"
restartPolicy: Never
Using partitionable devices
If your administrator has configured counter-based quota, your workload is charged by the device’s counter value (such as GPU memory) rather than device count. You submit workloads the same way as the ResourceClaimTemplate path above.
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
namespace: default
name: gpu-partition
spec:
spec:
devices:
requests:
- name: gpu
exactly:
deviceClassName: gpu.example.com
count: 1
---
apiVersion: batch/v1
kind: Job
metadata:
generateName: sample-dra-counter-job-
namespace: default
labels:
kueue.x-k8s.io/queue-name: user-queue
spec:
template:
spec:
containers:
- name: dummy-job
image: registry.k8s.io/e2e-test-images/agnhost:2.53
args: ["pause"]
resources:
claims:
- name: gpu
requests:
cpu: "1"
memory: "200Mi"
resourceClaims:
- name: gpu
resourceClaimTemplateName: gpu-partition
restartPolicy: Never
Using consumable capacity (shared devices)
Note
This feature requires theKueueDRAIntegrationConsumableCapacity feature gate,
which is disabled by default in v0.19.If your administrator has configured
capacity-based quota,
your workload is charged by the device’s capacity consumption (such as GPU
memory) rather than device count. You submit workloads using a
ResourceClaimTemplate with capacity.requests specifying how much capacity
you need:
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
namespace: default
name: gpu-share
spec:
spec:
devices:
requests:
- name: gpu
exactly:
deviceClassName: gpu.example.com
count: 1
capacity:
requests:
gpu.example.com/memory: "4Gi"
---
apiVersion: batch/v1
kind: Job
metadata:
generateName: sample-dra-capacity-job-
namespace: default
labels:
kueue.x-k8s.io/queue-name: user-queue
spec:
template:
spec:
containers:
- name: dummy-job
image: registry.k8s.io/e2e-test-images/agnhost:2.53
args: ["pause"]
resources:
claims:
- name: gpu
requests:
cpu: "1"
memory: "200Mi"
resourceClaims:
- name: gpu
resourceClaimTemplateName: gpu-share
restartPolicy: Never
If you omit capacity.requests, Kueue charges the device’s
RequestPolicy.Default or the full device capacity.
Using Topology-Aware Scheduling
Note
This feature requires theKueueDRADeviceFeasibility feature gate, which is
disabled by default in v0.20.If your administrator has set up
Topology-Aware Scheduling with DRA
for a Topology-Aware Scheduling (TAS)
queue, Kueue places each Pod only on nodes that can allocate the devices it
requests. You request devices the same way, with a ResourceClaimTemplate or an
extended resource, and add a topology annotation as for any TAS workload:
apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
namespace: default
name: single-gpu-tas
spec:
spec:
devices:
requests:
- name: gpu
exactly:
deviceClassName: gpu.example.com
---
apiVersion: batch/v1
kind: Job
metadata:
generateName: sample-dra-tas-job-
namespace: default
labels:
kueue.x-k8s.io/queue-name: dra-tas-user-queue
spec:
parallelism: 2
completions: 2
completionMode: Indexed
template:
metadata:
annotations:
kueue.x-k8s.io/podset-required-topology: "cloud.provider.com/topology-rack"
spec:
containers:
- name: dummy-job
image: registry.k8s.io/e2e-test-images/agnhost:2.53
args: ["pause"]
resources:
claims:
- name: gpu
requests:
cpu: "1"
memory: "200Mi"
resourceClaims:
- name: gpu
resourceClaimTemplateName: single-gpu-tas
restartPolicy: Never
If you are not sure which approach to use, ask your administrator.
2. Run the workload
You can run the workload with the following command.
For a ResourceClaimTemplate-based workload:
kubectl create -f https://kueue.sigs.k8s.io/examples/dra/sample-dra-rct-job.yaml
For a workload with a firstAvailable request:
kubectl create -f https://kueue.sigs.k8s.io/examples/dra/sample-dra-firstavailable-job.yaml
For an extended resource-based workload:
kubectl create -f https://kueue.sigs.k8s.io/examples/dra/sample-dra-extended-resource-job.yaml
For a workload in a Topology-Aware Scheduling queue:
kubectl create -f https://kueue.sigs.k8s.io/examples/dra/sample-dra-tas-job.yaml
If you submit the example more than once, kubectl reports that the
ResourceClaimTemplate single-gpu-tas already exists. The Job is still created.
Internally, Kueue will create a corresponding Workload for this Job.
3. (Optional) Monitor the status of the workload
You can see the Workload status with the following command:
kubectl -n default get workloads.kueue.x-k8s.io
To check whether the workload was admitted and see the DRA resource accounting:
kubectl -n default describe workload <workload-name>
Look at the Conditions section for admission status and the Events
section for details. If the workload was admitted, you can verify the
resources charged for quota in the
status.admission.podSetAssignments[].resourceUsage field:
kubectl -n default get workloads.kueue.x-k8s.io <workload-name> -o yaml
The Workload does not record which alternative a Pod received. Read it from the
Pod’s generated ResourceClaim, using the Job name kubectl create returned:
kubectl -n default get pods -l batch.kubernetes.io/job-name=<job-name> -o jsonpath='{.items[*].status.resourceClaimStatuses[*].resourceClaimName}'
Then, for each claim name printed:
kubectl -n default get resourceclaim <claim-name> -o jsonpath='{.status.allocation.devices.results[*].request}'
The output is similar to the following:
gpu/a100
Each value is <request>/<alternative>. The output is empty until the claim is
allocated.
The example container waits so that you can inspect the claim. When you finish, delete each example Job:
kubectl -n default delete job <job-name>
Once nothing uses a100-or-mig, delete the template:
kubectl -n default delete resourceclaimtemplate a100-or-mig
Troubleshooting
Workload not admitted
If the Workload stays in Pending state:
- Verify the
ClusterQueuehas quota for the DRA resource and it is not fully consumed by other workloads. - Run
kubectl -n default describe workload <workload-name>and look at the Events section for admission rejection reasons.
Workload pending with draNoFit
Run kubectl -n default describe workload <workload-name> and look at the
QuotaReserved condition in the Conditions section. If its reason is
TopologyPlacementFailed and its message includes draNoFit: N, no node in the
topology has the devices a single Pod requests. The devices may be in use by other
workloads, or tainted by an administrator if device taints are enabled. Kueue retries
when devices change, for example when another workload releases them or a taint is
removed. If no node can ever satisfy the request, reduce the number of devices each
Pod requests, or ask your administrator which nodes publish the DeviceClass you use.
If the message is Bypassed scheduling evaluation because an equivalent workload recently failed, Kueue skipped your workload because an equivalent one, with the same
requests, was just rejected. Look for another workload in the same queue whose message
includes draNoFit; it gives the reason. Both are retried when the devices change.
firstAvailable request not admitted
Run kubectl -n default describe workload <workload-name> and look at the
QuotaReserved condition. With reason DRAResourcesUnresolved (Inadmissible
if your administrator disabled UnadmittedWorkloadsObservability), the message
names the template and the field at fault. FirstAvailable device selection is not supported means KueueDRAIntegrationPrioritizedList is disabled, and
deviceClassName: Not found means the class is not in the mapping. Name a
mapped class, or ask your administrator to enable the gate or map the class and
restart the controller. For any other message, create a corrected template
under a new name and submit a new Job.
Pods pending after the workload is admitted
Kueue checks that a node can serve one Pod, not all the Pods it places there. If a
node has fewer free devices than the Pods placed on it request, some Pods stay
Pending until devices are released, for example when another workload finishes.
Double counting (extended resource path)
If quota usage shows double the expected value (e.g., 2 instead of 1 for
a single GPU), verify that KueueDRAIntegrationExtendedResource has not been
explicitly disabled. This gate is enabled by default since v0.19 and ensures
Kueue charges quota only once for extended resources backed by DRA, instead of
counting them as both a standard resource request and a DRA device.
Missing DeviceClass
For the extended resource path, the DeviceClass must exist before you submit
your workload. If it was created after your workload was rejected, the workload
may not be re-evaluated until another cluster event triggers requeuing.
Delete and re-create the workload to force re-evaluation.
For general troubleshooting, see the troubleshooting guide.
Feedback
Was this page helpful?
Glad to hear it! Please tell us how we can improve.
Sorry to hear that. Please tell us how we can improve.