Use Custom Preemption Configurations
This guide demonstrates how to configure and verify Configurable Preemptions in Kueue. You will learn how to:
- Enable the
ConfigurablePreemptionsandPrioritizePreemptorWorkloadsfeature gates. - Configure a dedicated Hero Workload Preemption configuration for an access-restricted
ClusterQueue. - Configure a Topology Defragmentation configuration for general-purpose workloads.
- Verify and observe preemption outcomes via eviction stats and status conditions.
Before you begin
Make sure the following conditions are met:
- A Kubernetes cluster running Kubernetes 1.30 or higher.
- Kueue v0.20.0 or higher installed.
- The
ConfigurablePreemptionsfeature gate enabled in the Kueue controller manager configuration. For hero workload scenarios, enablingPrioritizePreemptorWorkloadsis also recommended. (Note:TopologyAwareSchedulingis Beta and enabled by default since v0.14).
To enable the feature gates in your kueue-manager-config:
apiVersion: config.kueue.x-k8s.io/v1beta2
kind: Configuration
featureGates:
ConfigurablePreemptions: true
PrioritizePreemptorWorkloads: true
Configuration Scenarios
Rather than combining multiple concerns into a single configuration, it is recommended to define distinct PreemptionConfig resources tailored to specific queue purposes and operational privileges:
- Hero Workloads: Assigned to an access-restricted
ClusterQueuefor emergency or highest-priority jobs, granting elevated preemption privileges across queues. - Topology Defragmentation: Attached to general training
ClusterQueues, allowing large distributed jobs with feasible quota to preempt smaller fragmenting workloads when contiguous topology domains are unavailable.
Scenario 1: Hero Workload Preemption (Restricted Queue)
In this scenario, a dedicated, access-restricted ClusterQueue is established for top-priority distributed workloads (“hero workloads”). When hero workloads lack quota, they are permitted to preempt lower-priority workloads across the entire cohort hierarchy, while respecting protection labels on mission-critical queues.
1. Define the PreemptionConfig
apiVersion: kueue.x-k8s.io/v1alpha1
kind: PreemptionConfig
metadata:
name: "hero-workloads-preemption-config"
spec:
rules:
- name: "hero-preempt-lower-priority"
activationPolicy:
trigger: "Always"
candidateSelectors:
- scope: "AnyClusterQueue"
priority:
mode: "Base"
comparison: "LessThan"
clusterQueueSelector:
matchExpressions:
- key: "example.com/protection-tier"
operator: "NotIn"
values: ["mission-critical"]
How it works:
- Trigger:
Alwaysevaluates candidates unconditionally whenever preemption evaluation runs. This ensures the hero workload can preempt lower-priority workloads both to reclaim quota and to clear physical topology constraints under Topology-Aware Scheduling, preventing it from getting stuck on topology even after quota is satisfied. - Scope:
AnyClusterQueuesearches across all queues in the cluster. Because physical topology domains (e.g., racks or blocks) span the cluster and can be occupied by workloads from any queue, this allows the hero workload to unblock its topology requirements cluster-wide. - Elevated Privileges: Workloads in this queue can evict lower-priority workloads across queues even if those target workloads are running within their nominal quota (subject to overall cohort borrowing limits).
- Protection Guardrail:
clusterQueueSelectorensures that ClusterQueues labeledexample.com/protection-tier: mission-criticalare never selected for preemption.
2. Attach to the Restricted ClusterQueue
Attach the PreemptionConfig to your dedicated hero ClusterQueue:
apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata:
name: "hero-jobs-cq"
annotations:
kueue.x-k8s.io/preemption-config-name: "hero-workloads-preemption-config"
spec:
preemption:
reclaimWithinCohort: Never
withinClusterQueue: Never
# ... resource groups, flavors, and quotas ...
Important: Disabling Classical Preemption on Protected Queues
In Alpha, candidates selected byPreemptionConfig are merged with candidates selected by spec.preemption. If you configure spec.preemption.reclaimWithinCohort: LowerPriority, classical preemption will evaluate cohort candidates without checking the clusterQueueSelector in your PreemptionConfig. To ensure that protection labels are strictly honored, set reclaimWithinCohort: Never and withinClusterQueue: Never.Scenario 2: Topology Defragmentation (General Queue)
In this scenario, large distributed jobs require contiguous physical topology (such as full host blocks or racks under Topology-Aware Scheduling). A large workload may have feasible quota, but cannot be scheduled because smaller workloads are fragmenting the physical topology.
1. Define the PreemptionConfig
apiVersion: kueue.x-k8s.io/v1alpha1
kind: PreemptionConfig
metadata:
name: "topology-defrag-preemption-config"
spec:
rules:
- name: "preempt-lower-priority-within-cohort"
activationPolicy:
trigger: "InsufficientQuota"
candidateSelectors:
- scope: "WithinCohortTree"
priority:
mode: "Base"
comparison: "LessThan"
labelSelector:
matchExpressions:
- key: "example.com/workload-tier"
operator: "NotIn"
values: ["mission-critical"]
- name: "evict-smaller-jobs-for-topology"
activationPolicy:
trigger: "QuotaFeasibleAndInsufficientTopology"
candidateSelectors:
- scope: "AnyClusterQueue"
priority:
mode: "Base"
comparison: "LessThanOrEqual"
numericLabels:
- key: "example.com/node-count"
comparison: "LessThan"
labelSelector:
matchExpressions:
- key: "example.com/workload-tier"
operator: "NotIn"
values: ["mission-critical"]
How it works:
- Rule: preempt-lower-priority-within-cohort:
- Trigger:
InsufficientQuotaactivates when the incoming workload lacks sufficient quota to be admitted. - Scope:
WithinCohortTreerestricts quota reclamation to the cohort hierarchy (since borrowing outside the cohort tree is not permitted). - Protection Guardrail:
labelSelectorprevents evicting lower-priority workloads labeledexample.com/workload-tier: mission-critical.
- Trigger:
- Rule: evict-smaller-jobs-for-topology:
- Trigger:
QuotaFeasibleAndInsufficientTopologyactivates only when quota is already feasible for the incoming job under at least one eligible flavor assignment (after baseline preemption and any applicableInsufficientQuotarules), but placement is blocked by physical topology constraints.Note
QuotaFeasibleAndInsufficientTopologydoes not acquire missing quota—it only resolves topology fragmentation once quota feasibility has been satisfied. Combining this with anInsufficientQuotarule ensures workloads can first reclaim quota and then defragment topology. - Scope:
AnyClusterQueuesearches across all ClusterQueues in the cluster so topology can be unblocked across physical nodes regardless of cohort relationship. - Asymmetric Defragmentation:
numericLabelswithcomparison: LessThanensures that a larger workload (e.g.,example.com/node-count: 32) can preempt smaller workloads (e.g.,example.com/node-count: 4), but a 4-node workload cannot preempt a 32-node workload in return. OmittingfallbackValueensures unlabeled workloads are treated as incomparable and protected from eviction. - Protecting Mission-Critical Workloads: Without explicit exclusion, defragmentation rules could evict smaller mission-critical workloads. The
labelSelectorprevents evicting workloads labeledexample.com/workload-tier: mission-critical.
- Trigger:
2. Attach to the ClusterQueue
apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata:
name: "general-training-cq"
annotations:
kueue.x-k8s.io/preemption-config-name: "topology-defrag-preemption-config"
spec:
preemption:
reclaimWithinCohort: Never
withinClusterQueue: Never
# ... resource groups, flavors, and quotas ...
Disabling Classical Preemption for Label Protection
BecausePreemptionConfig now manages both quota acquisition and topology defragmentation while enforcing protection for example.com/workload-tier: mission-critical, set spec.preemption.reclaimWithinCohort: Never and spec.preemption.withinClusterQueue: Never. If classical preemption were left enabled as LowerPriority, it would evaluate cohort candidates without checking the labelSelector, potentially evicting lower-priority mission-critical workloads.Common Pitfalls
- Classical Preemption Bypassing Custom Guardrails: In Alpha, candidate sets from
spec.preemptionandPreemptionConfigare merged. If your custom rules protect specific workloads usinglabelSelectororclusterQueueSelector, classical preemption inspec.preemptionwill still evaluate and evict those workloads unless you setspec.preemption.reclaimWithinCohort: Neverandspec.preemption.withinClusterQueue: Never. - Preemption Flapping from Symmetric Rules: When defining rules across queues (with
WithinCohortTreeorAnyClusterQueue), ensure rules are strictly asymmetric (e.g., usingpriority.comparison: LessThanornumericLabels.comparison: LessThan) to avoid cascading preemptions where workloads repeatedly evict each other. - Label Propagation: Custom numeric labels or tier labels on Jobs are not copied to Kueue
Workloadresources unless added tointegrations.labelKeysToCopyin your Kueue Configuration. - Topology Defragmentation Requires Quota Feasibility: A workload blocked by both quota exhaustion and topology fragmentation cannot activate
QuotaFeasibleAndInsufficientTopologyuntil its quota requirement is satisfied by baseline preemption or anInsufficientQuotarule.
Verification & Observability
To provide visibility into why a workload was preempted when custom rules are used, Kueue tracks configurable preemption outcomes in workload status through eviction statistics and status conditions.
Eviction Scheduling Stats
Kueue records detailed eviction information in Workload.status.schedulingStats.evictions:
underlyingCause: Populated with the name of thePreemptionConfigthat triggered the preemption:status: schedulingStats: evictions: - count: 1 reason: Preempted underlyingCause: "hero-workloads-preemption-config"
Status Conditions
When a workload is preempted by a PreemptionConfig rule, Kueue sets two conditions in Workload.status.conditions. Both conditions share the exact same diagnostic message identifying the preemptor workload, the PreemptionConfig name, the rule name, and the selector index:
status:
conditions:
- type: Evicted
status: "True"
reason: Preempted
message: "Preempted by default/hero-job-xyz because of preemption config hero-workloads-preemption-config rule hero-preempt-lower-priority/0"
- type: Preempted
status: "True"
reason: ConfigurablePreemption
message: "Preempted by default/hero-job-xyz because of preemption config hero-workloads-preemption-config rule hero-preempt-lower-priority/0"
If multiple selectors within a rule triggered the candidate’s preemption, their indices are concatenated with commas (e.g., rule hero-preempt-lower-priority/0,1). If multiple rules contributed, they are separated with semicolons.
Metrics
Kueue exports Prometheus metrics broken down by queue and reason:
kueue_preempted_workloads_total{reason="ConfigurablePreemption"}: Counts workloads preempted byPreemptionConfigrules.kueue_admission_attempts_total{result="inadmissible"}: Counts failed admission attempts. InspectWorkload.status.conditionsto determine whether a workload was blocked by quota or topology.
Feedback
Was this page helpful?
Glad to hear it! Please tell us how we can improve.
Sorry to hear that. Please tell us how we can improve.