Configurable Preemptions
Cascading Preemption Risk
Configurable Preemptions is an advanced capability intended for power users and cluster administrators. Custom preemption rules carry inherent operational risks: if rules are misconfigured or symmetric (e.g., jobs can mutually preempt one another), they can trigger cascading preemptions and continuous job disruptions across the cluster.Configurable Preemptions introduces a declarative mechanism to define when preemption should occur and which workloads are eligible for eviction. It complements Kueue’s existing Classic Preemption and Fair Sharing algorithms by enabling policies for complex operational scenarios, such as:
- Topology Defragmentation: Allowing distributed workloads requiring specific physical topology domains (such as multi-node GPU or TPU training jobs under Topology-Aware Scheduling) to preempt smaller workloads that fragment the cluster, even when all workloads are within their nominal quotas.
- Mission-Critical “Hero” Workloads: Allowing dedicated, access-restricted queues with elevated preemption privileges to evict workloads across queues even when those workloads are within nominal quota (while remaining subject to configured cohort borrowing limits). When combined with the
PrioritizePreemptorWorkloadsfeature gate (Alpha in v0.20), hero jobs can effectively lock quota and gain admission without extra cluster-wide modifications. - Granular Priority & Label Rules: Evaluating candidates using either priorities or custom labels.
To use Configurable Preemptions, enable the ConfigurablePreemptions feature gate.
Architecture & API Overview
Configurable Preemption is driven by the cluster-scoped PreemptionConfig custom resource (kueue.x-k8s.io/v1alpha1). It contains a list of rules, each defining:
- Activation Policy (
activationPolicy): The trigger that activates the rule during a scheduling cycle. - Candidate Selectors (
candidateSelectors): One or more criteria that identify which running workloads may be considered for preemption.
Here is a complete example of a PreemptionConfig resource:
apiVersion: kueue.x-k8s.io/v1alpha1
kind: PreemptionConfig
metadata:
name: "topology-defragmentation"
spec:
rules:
- name: "evict-smaller-jobs-for-large-topology"
activationPolicy:
trigger: "QuotaFeasibleAndInsufficientTopology"
candidateSelectors:
- scope: "AnyClusterQueue"
priority:
mode: "Base"
comparison: "LessThanOrEqual"
numericLabels:
- key: "example.com/node-count"
comparison: "LessThan"
fallbackValue: 0
clusterQueueSelector:
matchExpressions:
- key: "example.com/protection-tier"
operator: "NotIn"
values: ["infrastructure"]
Activation Triggers
The activationPolicy.trigger field determines when a preemption rule becomes active during admission evaluation:
| Trigger | Description |
|---|---|
Always | Matching candidates are contributed unconditionally during preemption evaluation. |
InsufficientQuota | Matching candidates are contributed only if preempting baseline candidates does not yield sufficient quota to admit the incoming workload. |
QuotaFeasibleAndInsufficientTopology | Matching candidates are contributed only if quota is already feasible for the preemptor under at least one eligible flavor assignment (after baseline preemption and applicable InsufficientQuota candidates), but the workload cannot be admitted because no placement satisfies the physical topology requirements (see: Topology-Aware Scheduling). |
Candidate Selectors
Each rule specifies candidateSelectors to filter eligible preemption victims. Candidates resulting from multiple selectors within a rule (or across rules) are summed into a single deduplicated set. A candidate must satisfy all constraints specified within a selector:
1. Relational Scope (scope)
Required. The scope defines the relational boundary between the preemptor and candidate workloads:
WithinLocalQueue: Candidate must belong to the exact same LocalQueue as the preemptor (matching name and namespace).WithinClusterQueue: Candidate must belong to the exact same ClusterQueue as the preemptor.WithinParentCohort: Candidate belongs to a ClusterQueue sharing the immediate parent Cohort, or the preemptor’s own queue.WithinCohortTree: Candidate belongs to any ClusterQueue within the same root Cohort hierarchy, or the preemptor’s own queue.AnyClusterQueue: No relationship constraint; candidates can be selected from any ClusterQueue in the cluster.
2. Priority Constraints (priority)
Optional. Defines priority comparison criteria against the incoming preemptor workload. If specified, both mode and comparison are required:
mode(Required):Base: Compares raw priority values assigned inWorkload.spec.priority.Boosted: Compares effective priority values adjusted by priority boosting (see Priority Boosting).
comparison(Required):LessThan: Candidate priority < Preemptor priority.LessThanOrEqual: Candidate priority <= Preemptor priority.GreaterThan: Candidate priority > Preemptor priority.GreaterThanOrEqual: Candidate priority >= Preemptor priority.
3. Custom Numeric Labels (numericLabels)
Optional. A list of numeric label constraints allowing candidate filtering based on integer workload labels (e.g., number of GPUs/TPUs, slice index, or node count). Multiple numeric label constraints in a selector are joined using an AND rule (all constraints must be satisfied):
key(Required): The workload label key containing strings that can be parsed to integers as values.comparison: How the candidate’s label value compares to the preemptor’s label value (LessThan,LessThanOrEqual,GreaterThan,GreaterThanOrEqual).fallbackValue: Integer value assumed if a workload does not have the label or the value cannot be parsed. If omitted, workloads lacking the label are treated as incomparable and excluded.minValue/maxValue: Absolute lower and upper boundaries for the candidate’s label value.
Important
Custom labels from high-level jobs (e.g., Job, JobSet, RayCluster) are not automatically copied to the KueueWorkload resource unless their keys are listed in integrations.labelKeysToCopy in your Kueue Configuration. Ensure your custom numeric label keys are configured for copying.4. Label Selectors (labelSelector & clusterQueueSelector)
Optional. Standard Kubernetes label selectors:
labelSelector: Filters candidateWorkloadmetadata.clusterQueueSelector: Filters targetClusterQueuemetadata.
Referencing PreemptionConfig on a ClusterQueue
PreemptionConfig is attached to a ClusterQueue using the kueue.x-k8s.io/preemption-config-name annotation:
apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata:
name: "distributed-training-cq"
annotations:
kueue.x-k8s.io/preemption-config-name: "topology-defragmentation"
spec:
preemption:
reclaimWithinCohort: LowerPriority
withinClusterQueue: LowerPriority
# ... remaining ClusterQueue fields ...
Interaction with Existing Preemption
In Alpha,PreemptionConfig works alongside ClusterQueue.spec.preemption. Candidates from both mechanisms are evaluated and merged into a single candidate set before eviction.Candidate Merging and Ordering
During preemption evaluation in the scheduler, candidates from both mechanisms are gathered and combined:
- Dual Candidate Gathering:
- Classical / Fair Sharing candidates: Evaluated according to
ClusterQueue.spec.preemptionrules. Under Fair Sharing, candidates are prioritized according to Dominant Resource Sharing (DRS). - Configurable Preemptions candidates: Evaluated according to the active rules in the referenced
PreemptionConfig.
- Classical / Fair Sharing candidates: Evaluated according to
- Deduplication & Union: The scheduler combines candidates from both sources into a single set, deduplicating workloads by UID.
- Selective Control:
- To use only
PreemptionConfigrules and silence classical preemption, explicitly setspec.preemption.reclaimWithinCohort: Neverandspec.preemption.withinClusterQueue: Never. - To use only classical preemption, simply omit the
kueue.x-k8s.io/preemption-config-nameannotation.
- To use only
- Ordering & Evaluation: Once gathered, candidates are evaluated and sorted to satisfy the preemptor’s requirements. In Alpha, candidates from configurable rules are appended and integrated with standard preemption heuristics (such as prioritizing workloads already marked for eviction, cohort borrow status, priority, and admission recency).
Note on Beta Evolution
In Beta+,PreemptionConfig will achieve full feature parity with classical and fair sharing preemption. The strategies will become mutually exclusive via a formal API field on ClusterQueueSpec, and the Alpha annotation will be retired.Observability
To provide visibility into why a workload was preempted when custom rules are used, Kueue tracks configurable preemption outcomes in workload status:
1. Eviction Scheduling Stats
Kueue records detailed eviction information in Workload.status.schedulingStats.evictions:
reason: Set toPreempted.underlyingCause: Identifies the preemption config name that caused the preemption.
2. Status Conditions
Kueue also sets conditions in Workload.status.conditions:
status:
conditions:
- type: Evicted
status: "True"
reason: Preempted
message: "Preempted by default/hero-job-xyz because of preemption config hero-workloads-preemption-config rule hero-preempt-lower-priority/0"
- type: Preempted
status: "True"
reason: ConfigurablePreemption
message: "Preempted by default/hero-job-xyz because of preemption config hero-workloads-preemption-config rule hero-preempt-lower-priority/0"
Preventing Preemption Flapping
When authoring preemption rules across queues (especially with AnyClusterQueue or WithinCohortTree), ensure that rules are strictly asymmetric: if Workload A can preempt Workload B, Workload B must not be able to preempt Workload A in return.
Asymmetry can be guaranteed by:
- Requiring
priority.comparison: LessThan. - Enforcing
numericLabelswithcomparison: LessThanon job sizes. - Restricting preemption rights to dedicated high-priority queues.
What’s next?
- Follow the Use Custom Preemption Configurations guide for hands-on configuration steps and practical scenarios.
- Read Preemption to understand Classic Preemption and Fair Sharing algorithms.
- Read Topology-Aware Scheduling to see how physical network topology and defragmentation interact.
- Learn about Workload Priority Class to configure workload priorities.
Feedback
Was this page helpful?
Glad to hear it! Please tell us how we can improve.
Sorry to hear that. Please tell us how we can improve.