Run A RayCluster

Run a RayCluster with Kueue.

This page shows how to leverage Kueue’s scheduling and resource management capabilities when running RayCluster.

This guide is for batch users that have a basic understanding of Kueue. For more information, see Kueue’s overview.

Before you begin

  1. Make sure you are using Kueue v0.6.0 version or newer and KubeRay v1.1.0 or newer.

  2. Check Administer cluster quotas for details on the initial Kueue setup.

  3. See KubeRay Installation for installation and configuration details of KubeRay.

RayCluster definition

When running RayClusters on Kueue, take into consideration the following aspects:

a. Queue selection

The target local queue should be specified in the metadata.labels section of the RayCluster configuration.

metadata:
  labels:
    kueue.x-k8s.io/queue-name: user-queue

b. Configure the resource needs

The resource needs of the workload can be configured in the spec.

spec:
  headGroupSpec:
    template:
      spec:
        containers:
          - resources:
              requests:
                cpu: "1"
  workerGroupSpecs:
    - template:
        spec:
          containers:
            - resources:
                requests:
                  cpu: "1"

Note that a RayCluster will hold resource quotas while it exists. For optimal resource management, you should delete a RayCluster that is no longer in use.

c. Suspend control

Kueue controls the spec.suspend field of the RayCluster. When a RayCluster is admitted by Kueue, Kueue will unsuspend it by setting spec.suspend to false, regardless of its previous value.

d. Limitations

  • Limited Worker Groups: Because a Kueue workload can have a maximum of 18 PodSets, the maximum number of spec.workerGroupSpecs is 17

  • In-Tree Autoscaling Constraints: Autoscaling is only supported for elastic RayCluster objects. To enable in-tree autoscaling:

    1. Activate the ElasticJobsViaWorkloadSlices feature gate.

    2. Annotate the RayCluster object with:

      metadata:
        annotations:
          kueue.x-k8s.io/elastic-job: "true"
      
    3. Enable the Ray autoscaler of your RayCluster object by setting:

      spec:
        enableInTreeAutoscaling: true
      

Example RayCluster

The RayCluster looks like the following:

apiVersion: ray.io/v1
kind: RayCluster
metadata:
  name: raycluster-complete
  labels:
    kueue.x-k8s.io/queue-name: user-queue
spec:
  rayVersion: '2.55.1'
  headGroupSpec:
    serviceType: ClusterIP
    rayStartParams:
      dashboard-host: '0.0.0.0'
    template:
      spec:
        containers:
          - name: ray-head
            image: rayproject/ray:2.55.1
            resources:
              limits:
                cpu: "1"
                memory: "5Gi"
              requests:
                cpu: "1"
                memory: "5Gi"
            ports:
              - containerPort: 6379
                name: gcs
              - containerPort: 8265
                name: dashboard
              - containerPort: 10001
                name: client
            lifecycle:
              preStop:
                exec:
                  command: ["/bin/sh","-c","ray stop"]
            volumeMounts:
              - mountPath: /tmp/ray
                name: ray-logs
        volumes:
          - name: ray-logs
            emptyDir: {}
  workerGroupSpecs:
    - replicas: 1
      minReplicas: 1
      maxReplicas: 10
      groupName: small-group
      rayStartParams: {}
      template:
        spec:
          containers:
            - name: ray-worker
              image: rayproject/ray:2.55.1
              resources:
                limits:
                  cpu: "1"
                  memory: "1Gi"
                requests:
                  cpu: "1"
                  memory: "1Gi"
              lifecycle:
                preStop:
                  exec:
                    command: ["/bin/sh","-c","ray stop"]
              volumeMounts:
                - mountPath: /tmp/ray
                  name: ray-logs
          volumes:
            - name: ray-logs
              emptyDir: {}

You can submit a Ray Job using the CLI or log into the Ray Head and execute a job following this example with kind cluster.


Last modified September 30, 2026: Update main with the latest v0.20.0 (36c42c8df)