Dynamic Resource Allocation
Warning
In Kueue 0.18, the DRA feature gates were renamed to avoid conflicts with upstream Kubernetes feature gates:DynamicResourceAllocation is now KueueDRAIntegration,
and DRAExtendedResources is now KueueDRAIntegrationExtendedResource.Dynamic Resource Allocation
Dynamic Resource Allocation (DRA) is a Kubernetes API for requesting and managing hardware devices such as GPUs, FPGAs, and network adapters. Kueue can account for DRA devices in quota management through two paths:
ResourceClaimTemplate path: Pods explicitly reference a
ResourceClaimTemplatethat specifies a device request. Kueue maps eachDeviceClassreferenced by the claim to a logical resource name usingdeviceClassMappingsin the Kueue Configuration.Extended resource path: Pods request DRA devices using the traditional
resources.requestssyntax (e.g.,nvidia.com/gpu: 1). When the KubernetesDeviceClasshas anextendedResourceNamefield set (KEP-5004), the kube-scheduler automatically createsResourceClaimobjects from these requests. Kueue detects this and avoids double counting.
Note
DRA support in Kueue requires a Kubernetes cluster running version 1.34 or later where the DRA API (resource.k8s.io) is v1.Which path should I use? If your workloads already use resources.requests
for devices (e.g., nvidia.com/gpu: 1), use the extended resource path. If
your workloads explicitly create ResourceClaimTemplate objects, use the
ResourceClaimTemplate path.
How the ResourceClaimTemplate path works
When a Pod references a ResourceClaimTemplate, Kueue reads the
deviceClassName from the template’s exactly field and looks it up in
deviceClassMappings. With KueueDRAIntegrationPrioritizedList enabled it
reads a request’s firstAvailable alternatives as well; see the limitations
below for what that charges. This mapping tells Kueue which logical resource
name to charge quota against. The number of units charged is determined by the
count field in the device request (default 1).
Only the ExactCount allocation mode is supported. The
All allocation mode is not supported.
For setup instructions, see Set Up Dynamic Resource Allocation.
How the extended resource path works
When a Pod requests an extended resource backed by DRA (e.g.,
nvidia.com/gpu: 1), the kube-scheduler auto-creates a ResourceClaim.
Kueue detects the matching DeviceClass, uses extendedResourceName as the
quota key, and drops the auto-created claim from accounting. This prevents
quota from being charged for both the resources.requests entry and the
auto-created claim, which would double count the same device. No
deviceClassMappings configuration is needed; the mapping is discovered
from the DeviceClass automatically. A deviceClassMappings entry covering
that DeviceClass moves the charge to the mapping’s logical name.
This behavior is controlled by the KueueDRAIntegrationExtendedResource
feature gate, which is enabled by default since v0.19.
Note
The extended resource path additionally requires the KubernetesDRAExtendedResource feature gate on kube-apiserver and kube-scheduler
(beta in Kubernetes 1.36).Path separation
The two paths are independent:
- ResourceClaimTemplate path: uses
deviceClassMappingsconfiguration. - Extended resource path: uses auto-discovery from
DeviceClassobjects.
Do not configure the same DeviceClass in both paths for the same workload.
If overlap occurs, Kueue merges the resources using the deviceClassMappings
logical name as the quota key, which may result in incorrect quota accounting.
Quota accounting
DRA resources are tracked in ClusterQueue quotas just like CPU or memory.
The administrator includes the DRA resource name in coveredResources and
sets a nominalQuota. Kueue supports three quota accounting modes:
- Device count (default): Charges the
countvalue from the device request (default 1 when omitted). AClusterQueuewithexample.com/gpu: 8allows up to 8 concurrent device allocations. - Counter-based: Charges the device’s
consumesCountersvalue (e.g., GPU memory). See Counter-based quota. - Capacity-based: Charges the workload’s
capacity.requestsvalue rounded per the device’sRequestPolicy. See Capacity-based quota.
Admission and scheduling gap
There is a timing gap between Kueue admitting a workload (quota check) and the kube-scheduler allocating the actual device. Kueue does not know which specific device will be allocated — it only verifies that quota is available.
If the cluster state changes between these two steps (e.g., another system consumes the device), the scheduler may fail to allocate. The WaitForPodsReady feature provides a safety net by evicting workloads that fail to become ready within a configured timeout.
With Topology-Aware Scheduling, Kueue can also check before admission that a node has the devices a Pod needs. See Topology-Aware Scheduling with DRA.
Topology-Aware Scheduling with DRA
Note
KueueDRADeviceFeasibility is currently an alpha feature and is disabled by default.
You can enable it by editing the KueueDRADeviceFeasibility feature gate. Refer to the
Installation guide
for instructions on configuring feature gates. It requires KueueDRAIntegration,
TopologyAwareScheduling and TASNodeFeasibilityForAllLevels to be enabled as well.
KueueDRADeviceFeasibility adds a device check to
Topology-Aware Scheduling (TAS): TAS
places each Pod only on nodes that can allocate the devices it requests. Quota limits
how many devices a ClusterQueue admits, not where they are, and without this feature
TAS cannot tell which nodes have them. What goes wrong depends on how a Pod requests
its devices:
ResourceClaimTemplatepath: the devices are not in the Pod’s resource requests, so TAS places the Pod by its other resources alone. For example, aClusterQueuewith a quota of 8 GPUs on two nodes with 4 GPUs each admits a Pod whoseResourceClaimTemplaterequests 6 GPUs: the quota allows it, but no node has 6. The Pod staysPendingwhile the workload holds the quota.- Extended resource path: the Pod requests a resource such as
example.com/gpu, and TAS looks for it in each node’s allocatable, where a resource that only aDeviceClassprovides never appears, so the workload fits on no node; see the warning in When the check runs.
How the device check works
A workload is assigned a flavor with a
topologyName(a TAS flavor).Because a Pod’s
ResourceClaimobjects do not exist before admission, Kueue builds the claims each Pod will need: from itsResourceClaimTemplates, or from theDeviceClassfor an extended resource. For every node that the flavor and the Pod’s other scheduling constraints allow, Kueue tries to allocate these claims, using the same allocator as the kube-scheduler.Nodes where the allocation fails are dropped, and TAS places the Pods on the remaining nodes. For an extended resource, the check replaces TAS’s lookup in node allocatable, except on nodes that advertise the resource through a device plugin. This works for any topology, including one whose lowest level is not
kubernetes.io/hostname.When no node is left, the workload stays pending, and its
QuotaReservedcondition message counts the nodes rejected for devices asdraNoFit:couldn't assign flavors to pod set main: topology "dra-topology" doesn't allow to fit any of 1 pod(s). Total nodes: 2; excluded: draNoFit: 2Kueue checks the workload again when a
ResourceSliceorDeviceClasschanges, or when aResourceClaimreleases its devices.
When the check runs
| Workload | Device check |
|---|---|
Assigned a flavor with a topologyName, on the ResourceClaimTemplate or extended resource path | Runs |
Assigned a flavor without a topologyName | Does not run |
In a ClusterQueue with a MultiKueue admission check | Runs on the worker cluster, where topology is assigned, not on the manager |
In a ClusterQueue with a ProvisioningRequest admission check | Skipped on the first scheduling pass, which assigns no topology; runs on the second pass, after quota is reserved |
Warning
A workload on the extended resource path, such as one requestingexample.com/gpu: 1, requires the device check to be admitted to a TAS flavor. Without it,
TAS looks for the resource in each node’s allocatable, where a resource that only a
DeviceClass provides never appears, so the workload fits on no node:
excluded: resource "example.com/gpu": 2. Nodes that advertise the resource through a
device plugin are counted as before.Prerequisites
- A
Topologyand aResourceFlavorwithtopologyName, as described in Setup Topology-Aware Scheduling. - A DRA driver that publishes each node’s devices in
ResourceSliceobjects. - The
KueueDRADeviceFeasibilityfeature gate enabled in Kueue Configuration. The gates it requires are enabled by default; if one of them is disabled, Kueue does not start and logsconflicting feature gates detected.
Device taints
Note
KueueDRAIntegrationDeviceTaints is currently an alpha feature and is disabled by
default. It requires KueueDRADeviceFeasibility to be enabled as well.With this gate, the check skips devices with a NoSchedule or NoExecute
device taint
that the request does not tolerate, whether a DRA driver publishes the taint in a
ResourceSlice or an administrator applies it with a DeviceTaintRule. None taints
are ignored. A change to a DeviceTaintRule makes Kueue check rejected workloads again.
Kueue reads DeviceTaintRule objects only from Kubernetes 1.37 onwards, which serves them as
resource.k8s.io/v1; on earlier versions it ignores taints from rules. With this gate
disabled, it ignores all device taints. In both cases Kueue can admit a workload onto
tainted devices that the kube-scheduler then refuses, so enable this gate together
with KueueDRADeviceFeasibility.
Warning
The check runs only before admission. When aNoExecute taint is added to devices
that running Pods use, the Kubernetes eviction controller can delete those Pods, and
their replacements stay Pending, while Kueue keeps the workload admitted and its
quota reserved. Device taints do not trigger node replacement the way
TASReplaceNodeOnNodeTaints
does for node taints. To requeue such workloads, enable
WaitForPodsReady with a
recoveryTimeout; the check then keeps them pending until the taint is removed.Limitations of the check
- One Pod per node: the check asks whether a node can serve one Pod of the
PodSet, not how many. Kueue can place more Pods on a node than it has devices
for, and the Pods that do not get a device stay
Pending. - No release on preemption: devices held by workloads that Kueue would preempt are not freed in the check, so preemption cannot make a workload fit on devices.
- Kubernetes DRA feature gates are read from the Kueue process: the check follows the gates of the Kubernetes version Kueue is built with, Kubernetes 1.37 for Kueue v0.20, not the cluster’s. This matters only when objects carry the fields of a DRA feature that the kube-scheduler has disabled, for example when the feature is disabled on the kube-scheduler but not on the kube-apiserver, or disabled after objects already used it. Otherwise the kube-apiserver drops the fields of a disabled feature, so Kueue and the kube-scheduler see the same devices.
- Allocation time is not bounded: each check tries an allocation on every
node. A slow
DeviceClassCEL selector makes every scheduling cycle slower, rather than timing out. - Not every Kubernetes DRA feature is modeled: some, such as
DRADeviceBindingConditions, change what the kube-scheduler does but not what the check predicts (full list).
For setup instructions, see Use Topology-Aware Scheduling with DRA.
MultiKueue
DRA workloads are supported with MultiKueue,
except for firstAvailable requests; see the limitations below.
MultiKueue syncs the workload and its owning job to worker clusters, but
ResourceClaimTemplate and DeviceClass objects are not automatically
synced. These must be created on each worker cluster separately by the
cluster administrator.
Counter-based quota for partitionable devices
By default, Kueue tracks DRA quota by device count: each device request
charges count units regardless of the device’s capacity. This means a
small GPU partition and a full GPU both count as “1 device”, which does not
reflect the actual resource consumption.
Kueue can track quota using counter values published by DRA drivers
in ResourceSlice objects. This allows quota to reflect actual device
capacity (e.g., GPU memory) rather than device count.
This behavior is controlled by the KueueDRAIntegrationPartitionableDevices
feature gate, which is enabled by default since v0.19.
A DeviceClass uses either device-count quota (no sources configured) or
counter-based quota (with sources), not both. Kueue rejects configurations
that map the same DeviceClass to multiple resource names.
How it works
The administrator configures a
sourcesentry indeviceClassMappingsthat specifies which counter to track, which DRA driver to query, and a CEL expression to scope eligible devices.When a workload is submitted, Kueue reads the
consumesCountersfield from the matching devices inResourceSliceobjects to determine the actual counter charge.Kueue uses conservative charging: it takes the maximum
consumesCountersvalue across all matched devices and multiplies by the requestcount. This ensures quota is not undercharged when different devices consume different amounts.The
ClusterQueuequota is set in counter units (e.g.,800Gifor GPU memory) instead of device count.
Prerequisites
- Kubernetes 1.35 or later with the
DRAPartitionableDevicesfeature gate enabled (beta in Kubernetes 1.36). - A DRA driver that publishes
consumesCounterson devices inResourceSliceobjects.
For setup instructions, see Set Up Dynamic Resource Allocation.
Counter-based vs capacity-based quota
Both modes track quota by actual resource consumption rather than device count, but they serve different device types:
| Counter-based (PD) | Capacity-based (CC) | |
|---|---|---|
| Device type | Partitioned devices (e.g., NVIDIA MIG) | Shared devices (e.g., GPU time-slicing, MPS) |
| Charge source | Device’s consumesCounters | Workload’s capacity.requests |
| Who decides consumption | Driver (fixed per partition) | User (variable per workload) |
| Upstream K8s feature | KEP-4815 (DRAPartitionableDevices) | KEP-5075 (DRAConsumableCapacity) |
If your GPUs use hardware partitioning (MIG), use counter-based quota. If your GPUs allow software-level sharing where workloads request variable amounts of capacity, use capacity-based quota.
A cluster can use both modes simultaneously with different DeviceClasses using the same DRA driver. One DeviceClass with counter sources for partitioned devices and another with capacity sources for shared devices. Counter and capacity sources cannot be mixed within the same DeviceClass mapping.
Capacity-based quota for shared devices (consumable capacity)
Some devices allow multiple workloads to share them simultaneously using
software-level sharing mechanisms such as GPU time-slicing or MPS. These
devices publish a Capacity field on each device in ResourceSlice objects
(defined by KEP-5075)
instead of using consumesCounters. Workloads specify how much capacity they
need via capacity.requests on the device request.
Kueue can track quota using these capacity dimensions so that the total consumed capacity across all sharing workloads does not exceed the device’s published capacity.
This behavior is controlled by the KueueDRAIntegrationConsumableCapacity
feature gate (Alpha, disabled by default in v0.19).
A DeviceClass uses either device-count quota (no sources), counter-based
quota (with counter sources), or capacity-based quota (with capacity
sources). Counter and capacity sources cannot be mixed in the same mapping.
How it works
The administrator configures a
capacitysource entry indeviceClassMappingsthat specifies which capacity dimension to track, which DRA driver to query, and a CEL expression to scope eligible devices.When a workload is submitted, Kueue reads the workload’s
capacity.requestsfrom theExactDeviceRequestfor the configured dimension. Ifcapacity.requestsis omitted, Kueue uses the device’sRequestPolicy.Defaultor the fullCapacity.Valueas the charge.Kueue rounds the request per the device’s
RequestPolicy(ValidValuesorValidRangewithStep) to prevent quota gaming where a small request consumes more actual capacity after rounding by the kube-scheduler.For each matched device, Kueue computes the charge independently using the device’s own Default and policy, then takes the maximum across all devices. This ensures quota is never undercharged even if the
deviceSelectormatches heterogeneous devices.The
ClusterQueuequota is set in capacity units (e.g.,800Gifor GPU memory) instead of device count.
Prerequisites
- Kubernetes 1.36 or later with the
DRAConsumableCapacityfeature gate enabled (beta, enabled by default in Kubernetes 1.36). - A DRA driver that publishes
CapacityandAllowMultipleAllocationson devices inResourceSliceobjects. - The
KueueDRAIntegrationConsumableCapacityfeature gate enabled in Kueue Configuration.
For setup instructions, see Set Up Dynamic Resource Allocation.
Limitations
The following limitations apply:
- ResourceClaimTemplates only: Only
ResourceClaimTemplatereferences are supported. DirectResourceClaimreferences in the Pod spec are not supported and will result in inadmissible workloads. - ExactCount allocation mode only: the
Allallocation mode is not supported, in anexactlyrequest or in an alternative of afirstAvailableone. Kueue reads afirstAvailablerequest only when theKueueDRAIntegrationPrioritizedListfeature gate is enabled. That gate is alpha and off by default; see the note below for what it covers. - No device constraints or config: Device
constraints(MatchAttribute) and per-requestconfigare not supported. - No AdminAccess: Device requests with
adminAccess: trueare not supported. - TAS does not see devices: without
Topology-Aware Scheduling with DRA, TAS may
place a Pod on a node without the devices its
ResourceClaimTemplaterequests, and a workload that requests an extended resource only aDeviceClassprovides fits on no node. - Device taints do not change quota: A tainted device is charged like any other. The per-node check can honor taints; see Device taints.
firstAvailablerequests are charged, within limits: This support is experimental; do not enable it in production. WithKueueDRAIntegrationPrioritizedListenabled, afirstAvailablerequest is charged once, the count every alternative asks for, which is what the scheduler allocates whichever alternative it picks. Every alternative of a request has to ask for the same count and map to the same logical resource; a request whose alternatives differ in count is refused, and so is an alternative on a mapping with acounterorcapacitysource. An alternative that setscapacityon the subrequest is charged its declared count like any other. WithoutKueueDRADeviceFeasibility, Kueue does not check that any alternative can be satisfied by the cluster, so a request whose alternatives are all infeasible holds its quota until the Workload is evicted, for example by WaitForPodsReady where it is configured; with it, such a Workload stays pending instead. The charge lands on the one ResourceFlavor the PodSet is assigned, and the Pods carry that flavor’s node labels, so keep every alternative’s devices behind the same flavors; with a flavor per device model, only the alternative with devices on the assigned flavor can run. MultiKueue does not supportfirstAvailablerequests: a manager and a worker may resolve different templates, and nothing refuses such a Workload before dispatch yet.
Feedback
Was this page helpful?
Glad to hear it! Please tell us how we can improve.
Sorry to hear that. Please tell us how we can improve.