Configurable Preemptions

Declarative, rule-based preemption policies for custom candidate selection.
Feature state alpha since Kueue v0.20

Configurable Preemptions introduces a declarative mechanism to define when preemption should occur and which workloads are eligible for eviction. It complements Kueue’s existing Classic Preemption and Fair Sharing algorithms by enabling policies for complex operational scenarios, such as:

  • Topology Defragmentation: Allowing distributed workloads requiring specific physical topology domains (such as multi-node GPU or TPU training jobs under Topology-Aware Scheduling) to preempt smaller workloads that fragment the cluster, even when all workloads are within their nominal quotas.
  • Mission-Critical “Hero” Workloads: Allowing dedicated, access-restricted queues with elevated preemption privileges to evict workloads across queues even when those workloads are within nominal quota (while remaining subject to configured cohort borrowing limits). When combined with the PrioritizePreemptorWorkloads feature gate (Alpha in v0.20), hero jobs can effectively lock quota and gain admission without extra cluster-wide modifications.
  • Granular Priority & Label Rules: Evaluating candidates using either priorities or custom labels.

To use Configurable Preemptions, enable the ConfigurablePreemptions feature gate.

Architecture & API Overview

Configurable Preemption is driven by the cluster-scoped PreemptionConfig custom resource (kueue.x-k8s.io/v1alpha1). It contains a list of rules, each defining:

  1. Activation Policy (activationPolicy): The trigger that activates the rule during a scheduling cycle.
  2. Candidate Selectors (candidateSelectors): One or more criteria that identify which running workloads may be considered for preemption.

Here is a complete example of a PreemptionConfig resource:

apiVersion: kueue.x-k8s.io/v1alpha1
kind: PreemptionConfig
metadata:
  name: "topology-defragmentation"
spec:
  rules:
  - name: "evict-smaller-jobs-for-large-topology"
    activationPolicy:
      trigger: "QuotaFeasibleAndInsufficientTopology"
    candidateSelectors:
    - scope: "AnyClusterQueue"
      priority:
        mode: "Base"
        comparison: "LessThanOrEqual"
      numericLabels:
      - key: "example.com/node-count"
        comparison: "LessThan"
        fallbackValue: 0
      clusterQueueSelector:
        matchExpressions:
        - key: "example.com/protection-tier"
          operator: "NotIn"
          values: ["infrastructure"]

Activation Triggers

The activationPolicy.trigger field determines when a preemption rule becomes active during admission evaluation:

TriggerDescription
AlwaysMatching candidates are contributed unconditionally during preemption evaluation.
InsufficientQuotaMatching candidates are contributed only if preempting baseline candidates does not yield sufficient quota to admit the incoming workload.
QuotaFeasibleAndInsufficientTopologyMatching candidates are contributed only if quota is already feasible for the preemptor under at least one eligible flavor assignment (after baseline preemption and applicable InsufficientQuota candidates), but the workload cannot be admitted because no placement satisfies the physical topology requirements (see: Topology-Aware Scheduling).

Candidate Selectors

Each rule specifies candidateSelectors to filter eligible preemption victims. Candidates resulting from multiple selectors within a rule (or across rules) are summed into a single deduplicated set. A candidate must satisfy all constraints specified within a selector:

1. Relational Scope (scope)

Required. The scope defines the relational boundary between the preemptor and candidate workloads:

  • WithinLocalQueue: Candidate must belong to the exact same LocalQueue as the preemptor (matching name and namespace).
  • WithinClusterQueue: Candidate must belong to the exact same ClusterQueue as the preemptor.
  • WithinParentCohort: Candidate belongs to a ClusterQueue sharing the immediate parent Cohort, or the preemptor’s own queue.
  • WithinCohortTree: Candidate belongs to any ClusterQueue within the same root Cohort hierarchy, or the preemptor’s own queue.
  • AnyClusterQueue: No relationship constraint; candidates can be selected from any ClusterQueue in the cluster.

2. Priority Constraints (priority)

Optional. Defines priority comparison criteria against the incoming preemptor workload. If specified, both mode and comparison are required:

  • mode (Required):
    • Base: Compares raw priority values assigned in Workload.spec.priority.
    • Boosted: Compares effective priority values adjusted by priority boosting (see Priority Boosting).
  • comparison (Required):
    • LessThan: Candidate priority < Preemptor priority.
    • LessThanOrEqual: Candidate priority <= Preemptor priority.
    • GreaterThan: Candidate priority > Preemptor priority.
    • GreaterThanOrEqual: Candidate priority >= Preemptor priority.

3. Custom Numeric Labels (numericLabels)

Optional. A list of numeric label constraints allowing candidate filtering based on integer workload labels (e.g., number of GPUs/TPUs, slice index, or node count). Multiple numeric label constraints in a selector are joined using an AND rule (all constraints must be satisfied):

  • key (Required): The workload label key containing strings that can be parsed to integers as values.
  • comparison: How the candidate’s label value compares to the preemptor’s label value (LessThan, LessThanOrEqual, GreaterThan, GreaterThanOrEqual).
  • fallbackValue: Integer value assumed if a workload does not have the label or the value cannot be parsed. If omitted, workloads lacking the label are treated as incomparable and excluded.
  • minValue / maxValue: Absolute lower and upper boundaries for the candidate’s label value.

4. Label Selectors (labelSelector & clusterQueueSelector)

Optional. Standard Kubernetes label selectors:

  • labelSelector: Filters candidate Workload metadata.
  • clusterQueueSelector: Filters target ClusterQueue metadata.

Referencing PreemptionConfig on a ClusterQueue

PreemptionConfig is attached to a ClusterQueue using the kueue.x-k8s.io/preemption-config-name annotation:

apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata:
  name: "distributed-training-cq"
  annotations:
    kueue.x-k8s.io/preemption-config-name: "topology-defragmentation"
spec:
  preemption:
    reclaimWithinCohort: LowerPriority
    withinClusterQueue: LowerPriority
  # ... remaining ClusterQueue fields ...

Candidate Merging and Ordering

During preemption evaluation in the scheduler, candidates from both mechanisms are gathered and combined:

  1. Dual Candidate Gathering:
    • Classical / Fair Sharing candidates: Evaluated according to ClusterQueue.spec.preemption rules. Under Fair Sharing, candidates are prioritized according to Dominant Resource Sharing (DRS).
    • Configurable Preemptions candidates: Evaluated according to the active rules in the referenced PreemptionConfig.
  2. Deduplication & Union: The scheduler combines candidates from both sources into a single set, deduplicating workloads by UID.
  3. Selective Control:
    • To use only PreemptionConfig rules and silence classical preemption, explicitly set spec.preemption.reclaimWithinCohort: Never and spec.preemption.withinClusterQueue: Never.
    • To use only classical preemption, simply omit the kueue.x-k8s.io/preemption-config-name annotation.
  4. Ordering & Evaluation: Once gathered, candidates are evaluated and sorted to satisfy the preemptor’s requirements. In Alpha, candidates from configurable rules are appended and integrated with standard preemption heuristics (such as prioritizing workloads already marked for eviction, cohort borrow status, priority, and admission recency).

Observability

To provide visibility into why a workload was preempted when custom rules are used, Kueue tracks configurable preemption outcomes in workload status:

1. Eviction Scheduling Stats

Kueue records detailed eviction information in Workload.status.schedulingStats.evictions:

  • reason: Set to Preempted.
  • underlyingCause: Identifies the preemption config name that caused the preemption.

2. Status Conditions

Kueue also sets conditions in Workload.status.conditions:

status:
  conditions:
  - type: Evicted
    status: "True"
    reason: Preempted
    message: "Preempted by default/hero-job-xyz because of preemption config hero-workloads-preemption-config rule hero-preempt-lower-priority/0"
  - type: Preempted
    status: "True"
    reason: ConfigurablePreemption
    message: "Preempted by default/hero-job-xyz because of preemption config hero-workloads-preemption-config rule hero-preempt-lower-priority/0"

Preventing Preemption Flapping

When authoring preemption rules across queues (especially with AnyClusterQueue or WithinCohortTree), ensure that rules are strictly asymmetric: if Workload A can preempt Workload B, Workload B must not be able to preempt Workload A in return.

Asymmetry can be guaranteed by:

  • Requiring priority.comparison: LessThan.
  • Enforcing numericLabels with comparison: LessThan on job sizes.
  • Restricting preemption rights to dedicated high-priority queues.

What’s next?