Use Custom Preemption Configurations

Set up and verify declarative preemption configurations for hero workloads and topology defragmentation using PreemptionConfig.
Feature state alpha since Kueue v0.20

This guide demonstrates how to configure and verify Configurable Preemptions in Kueue. You will learn how to:

  1. Enable the ConfigurablePreemptions and PrioritizePreemptorWorkloads feature gates.
  2. Configure a dedicated Hero Workload Preemption configuration for an access-restricted ClusterQueue.
  3. Configure a Topology Defragmentation configuration for general-purpose workloads.
  4. Verify and observe preemption outcomes via eviction stats and status conditions.

Before you begin

Make sure the following conditions are met:

  • A Kubernetes cluster running Kubernetes 1.30 or higher.
  • Kueue v0.20.0 or higher installed.
  • The ConfigurablePreemptions feature gate enabled in the Kueue controller manager configuration. For hero workload scenarios, enabling PrioritizePreemptorWorkloads is also recommended. (Note: TopologyAwareScheduling is Beta and enabled by default since v0.14).

To enable the feature gates in your kueue-manager-config:

apiVersion: config.kueue.x-k8s.io/v1beta2
kind: Configuration
featureGates:
  ConfigurablePreemptions: true
  PrioritizePreemptorWorkloads: true

Configuration Scenarios

Rather than combining multiple concerns into a single configuration, it is recommended to define distinct PreemptionConfig resources tailored to specific queue purposes and operational privileges:

  1. Hero Workloads: Assigned to an access-restricted ClusterQueue for emergency or highest-priority jobs, granting elevated preemption privileges across queues.
  2. Topology Defragmentation: Attached to general training ClusterQueues, allowing large distributed jobs with feasible quota to preempt smaller fragmenting workloads when contiguous topology domains are unavailable.

Scenario 1: Hero Workload Preemption (Restricted Queue)

In this scenario, a dedicated, access-restricted ClusterQueue is established for top-priority distributed workloads (“hero workloads”). When hero workloads lack quota, they are permitted to preempt lower-priority workloads across the entire cohort hierarchy, while respecting protection labels on mission-critical queues.

1. Define the PreemptionConfig

apiVersion: kueue.x-k8s.io/v1alpha1
kind: PreemptionConfig
metadata:
  name: "hero-workloads-preemption-config"
spec:
  rules:
  - name: "hero-preempt-lower-priority"
    activationPolicy:
      trigger: "Always"
    candidateSelectors:
    - scope: "AnyClusterQueue"
      priority:
        mode: "Base"
        comparison: "LessThan"
      clusterQueueSelector:
        matchExpressions:
        - key: "example.com/protection-tier"
          operator: "NotIn"
          values: ["mission-critical"]

How it works:

  • Trigger: Always evaluates candidates unconditionally whenever preemption evaluation runs. This ensures the hero workload can preempt lower-priority workloads both to reclaim quota and to clear physical topology constraints under Topology-Aware Scheduling, preventing it from getting stuck on topology even after quota is satisfied.
  • Scope: AnyClusterQueue searches across all queues in the cluster. Because physical topology domains (e.g., racks or blocks) span the cluster and can be occupied by workloads from any queue, this allows the hero workload to unblock its topology requirements cluster-wide.
  • Elevated Privileges: Workloads in this queue can evict lower-priority workloads across queues even if those target workloads are running within their nominal quota (subject to overall cohort borrowing limits).
  • Protection Guardrail: clusterQueueSelector ensures that ClusterQueues labeled example.com/protection-tier: mission-critical are never selected for preemption.

2. Attach to the Restricted ClusterQueue

Attach the PreemptionConfig to your dedicated hero ClusterQueue:

apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata:
  name: "hero-jobs-cq"
  annotations:
    kueue.x-k8s.io/preemption-config-name: "hero-workloads-preemption-config"
spec:
  preemption:
    reclaimWithinCohort: Never
    withinClusterQueue: Never
  # ... resource groups, flavors, and quotas ...

Scenario 2: Topology Defragmentation (General Queue)

In this scenario, large distributed jobs require contiguous physical topology (such as full host blocks or racks under Topology-Aware Scheduling). A large workload may have feasible quota, but cannot be scheduled because smaller workloads are fragmenting the physical topology.

1. Define the PreemptionConfig

apiVersion: kueue.x-k8s.io/v1alpha1
kind: PreemptionConfig
metadata:
  name: "topology-defrag-preemption-config"
spec:
  rules:
  - name: "preempt-lower-priority-within-cohort"
    activationPolicy:
      trigger: "InsufficientQuota"
    candidateSelectors:
    - scope: "WithinCohortTree"
      priority:
        mode: "Base"
        comparison: "LessThan"
      labelSelector:
        matchExpressions:
        - key: "example.com/workload-tier"
          operator: "NotIn"
          values: ["mission-critical"]
  - name: "evict-smaller-jobs-for-topology"
    activationPolicy:
      trigger: "QuotaFeasibleAndInsufficientTopology"
    candidateSelectors:
    - scope: "AnyClusterQueue"
      priority:
        mode: "Base"
        comparison: "LessThanOrEqual"
      numericLabels:
      - key: "example.com/node-count"
        comparison: "LessThan"
      labelSelector:
        matchExpressions:
        - key: "example.com/workload-tier"
          operator: "NotIn"
          values: ["mission-critical"]

How it works:

  • Rule: preempt-lower-priority-within-cohort:
    • Trigger: InsufficientQuota activates when the incoming workload lacks sufficient quota to be admitted.
    • Scope: WithinCohortTree restricts quota reclamation to the cohort hierarchy (since borrowing outside the cohort tree is not permitted).
    • Protection Guardrail: labelSelector prevents evicting lower-priority workloads labeled example.com/workload-tier: mission-critical.
  • Rule: evict-smaller-jobs-for-topology:
    • Trigger: QuotaFeasibleAndInsufficientTopology activates only when quota is already feasible for the incoming job under at least one eligible flavor assignment (after baseline preemption and any applicable InsufficientQuota rules), but placement is blocked by physical topology constraints.
    • Scope: AnyClusterQueue searches across all ClusterQueues in the cluster so topology can be unblocked across physical nodes regardless of cohort relationship.
    • Asymmetric Defragmentation: numericLabels with comparison: LessThan ensures that a larger workload (e.g., example.com/node-count: 32) can preempt smaller workloads (e.g., example.com/node-count: 4), but a 4-node workload cannot preempt a 32-node workload in return. Omitting fallbackValue ensures unlabeled workloads are treated as incomparable and protected from eviction.
    • Protecting Mission-Critical Workloads: Without explicit exclusion, defragmentation rules could evict smaller mission-critical workloads. The labelSelector prevents evicting workloads labeled example.com/workload-tier: mission-critical.

2. Attach to the ClusterQueue

apiVersion: kueue.x-k8s.io/v1beta2
kind: ClusterQueue
metadata:
  name: "general-training-cq"
  annotations:
    kueue.x-k8s.io/preemption-config-name: "topology-defrag-preemption-config"
spec:
  preemption:
    reclaimWithinCohort: Never
    withinClusterQueue: Never
  # ... resource groups, flavors, and quotas ...

Common Pitfalls

  • Classical Preemption Bypassing Custom Guardrails: In Alpha, candidate sets from spec.preemption and PreemptionConfig are merged. If your custom rules protect specific workloads using labelSelector or clusterQueueSelector, classical preemption in spec.preemption will still evaluate and evict those workloads unless you set spec.preemption.reclaimWithinCohort: Never and spec.preemption.withinClusterQueue: Never.
  • Preemption Flapping from Symmetric Rules: When defining rules across queues (with WithinCohortTree or AnyClusterQueue), ensure rules are strictly asymmetric (e.g., using priority.comparison: LessThan or numericLabels.comparison: LessThan) to avoid cascading preemptions where workloads repeatedly evict each other.
  • Label Propagation: Custom numeric labels or tier labels on Jobs are not copied to Kueue Workload resources unless added to integrations.labelKeysToCopy in your Kueue Configuration.
  • Topology Defragmentation Requires Quota Feasibility: A workload blocked by both quota exhaustion and topology fragmentation cannot activate QuotaFeasibleAndInsufficientTopology until its quota requirement is satisfied by baseline preemption or an InsufficientQuota rule.

Verification & Observability

To provide visibility into why a workload was preempted when custom rules are used, Kueue tracks configurable preemption outcomes in workload status through eviction statistics and status conditions.

Eviction Scheduling Stats

Kueue records detailed eviction information in Workload.status.schedulingStats.evictions:

  • underlyingCause: Populated with the name of the PreemptionConfig that triggered the preemption:
    status:
      schedulingStats:
        evictions:
        - count: 1
          reason: Preempted
          underlyingCause: "hero-workloads-preemption-config"
    

Status Conditions

When a workload is preempted by a PreemptionConfig rule, Kueue sets two conditions in Workload.status.conditions. Both conditions share the exact same diagnostic message identifying the preemptor workload, the PreemptionConfig name, the rule name, and the selector index:

status:
  conditions:
  - type: Evicted
    status: "True"
    reason: Preempted
    message: "Preempted by default/hero-job-xyz because of preemption config hero-workloads-preemption-config rule hero-preempt-lower-priority/0"
  - type: Preempted
    status: "True"
    reason: ConfigurablePreemption
    message: "Preempted by default/hero-job-xyz because of preemption config hero-workloads-preemption-config rule hero-preempt-lower-priority/0"

If multiple selectors within a rule triggered the candidate’s preemption, their indices are concatenated with commas (e.g., rule hero-preempt-lower-priority/0,1). If multiple rules contributed, they are separated with semicolons.

Metrics

Kueue exports Prometheus metrics broken down by queue and reason:

  • kueue_preempted_workloads_total{reason="ConfigurablePreemption"}: Counts workloads preempted by PreemptionConfig rules.
  • kueue_admission_attempts_total{result="inadmissible"}: Counts failed admission attempts. Inspect Workload.status.conditions to determine whether a workload was blocked by quota or topology.