Run A RayService

Run a RayService with Kueue.

This page shows how to leverage Kueue’s scheduling and resource management capabilities when running RayService.

Kueue manages the RayService as a top-level job, similar to how Kueue manages the RayJob.

This guide is for serving users that have a basic understanding of Kueue. For more information, see Kueue’s overview.

Before you begin

  1. Make sure you are using Kueue v0.17.0 or newer and KubeRay v1.3.0 or newer.

  2. Check Administer cluster quotas for details on the initial Kueue setup.

  3. See KubeRay Installation for installation and configuration details of KubeRay.

RayService definition

When running RayService on Kueue, take into consideration the following aspects:

a. Queue selection

The target local queue should be specified in the metadata.labels section of the RayService configuration, and this label will be propagated to its RayCluster.

metadata:
  labels:
    kueue.x-k8s.io/queue-name: user-queue

b. Configure the resource needs

The resource needs of the workload can be configured in the spec.rayClusterConfig.

spec:
  rayClusterConfig:
    headGroupSpec:
    template:
      spec:
        containers:
          - resources:
              requests:
                cpu: "1"
    workerGroupSpecs:
    - template:
        spec:
          containers:
            - resources:
                requests:
                  cpu: "1"

c. Suspend control

Kueue controls the spec.rayClusterConfig.suspend field of the RayService. When a RayService is admitted by Kueue, Kueue will unsuspend it by setting spec.rayClusterConfig.suspend to false, regardless of its previous value.

d. Limitations

  • Limited Worker Groups: Because a Kueue workload can have a maximum of 18 PodSets, the maximum number of spec.rayClusterConfig.workerGroupSpecs is 17.

Autoscaling (a.k.a InTreeAutoscaling)

Ray Autoscaling can automatically add or remove Ray worker pods based on resource demand.

This feature is supported for RayService starting in Kueue v0.19.0.

Autoscaling is only supported for elastic RayService objects.

How to enable autoscaling in RayService

  1. Enable the feature gate for Elastic Workloads (Workload Slices):

    ElasticJobsViaWorkloadSlices: true
    
  2. Add the workload slicing annotation to the RayService object:

    metadata:
      annotations:
        kueue.x-k8s.io/elastic-job: "true"
    
  3. Enable the Ray autoscaler in the RayService object:

    spec:
      rayClusterConfig:
        enableInTreeAutoscaling: true
    

Rolling upgrades limitation

Kueue’s Workload Slices feature currently manages quota for a single active cluster. Upgrade strategies that provision a secondary surge cluster (spec.upgradeStrategy.type: NewCluster or NewClusterWithIncrementalUpgrade) are not currently supported when workload slicing is enabled because pending cluster pods will remain gated. To use workload slicing with autoscaling, use spec.upgradeStrategy.type: None or apply updates in-place.

Example RayService

The RayService looks like the following:

apiVersion: ray.io/v1
kind: RayService
metadata:
  name: test-rayservice
  namespace: default
  labels:
    kueue.x-k8s.io/queue-name: user-queue
spec:
  upgradeStrategy:
    type: None
  serveConfigV2: |
    applications:
      - name: fruit_app
        import_path: fruit.deployment_graph
        route_prefix: /fruit
        runtime_env:
          working_dir: "https://github.com/ray-project/test_dag/archive/78b4a5da38796123d9f9ffff59bab2792a043e95.zip"
        deployments:
          - name: MangoStand
            num_replicas: 2
            max_replicas_per_node: 1
            user_config:
              price: 3
            ray_actor_options:
              num_cpus: 0.1
          - name: OrangeStand
            num_replicas: 1
            user_config:
              price: 2
            ray_actor_options:
              num_cpus: 0.1
          - name: PearStand
            num_replicas: 1
            user_config:
              price: 1
            ray_actor_options:
              num_cpus: 0.1
          - name: FruitMarket
            num_replicas: 1
            ray_actor_options:
              num_cpus: 0.1
      - name: math_app
        import_path: conditional_dag.serve_dag
        route_prefix: /calc
        runtime_env:
          working_dir: "https://github.com/ray-project/test_dag/archive/78b4a5da38796123d9f9ffff59bab2792a043e95.zip"
        deployments:
          - name: Adder
            num_replicas: 1
            user_config:
              increment: 3
            ray_actor_options:
              num_cpus: 0.1
          - name: Multiplier
            num_replicas: 1
            user_config:
              factor: 5
            ray_actor_options:
              num_cpus: 0.1
          - name: Router
            num_replicas: 1
  rayClusterConfig:
    rayVersion: '2.55.1'
    headGroupSpec:
      rayStartParams: {}
      template:
        spec:
          containers:
          - name: ray-head
            image: rayproject/ray:2.55.1
            resources:
              limits:
                cpu: "2"
                memory: "5Gi"
              requests:
                cpu: "2"
                memory: "2Gi"
    workerGroupSpecs:
    - replicas: 1
      minReplicas: 1
      maxReplicas: 5
      groupName: small-group
      rayStartParams: {}
      template:
        spec:
          containers:
          - name: ray-worker
            image: rayproject/ray:2.55.1
            resources:
              limits:
                cpu: "1"
                memory: "2Gi"
              requests:
                cpu: "500m"
                memory: "2Gi"

Last modified September 30, 2026: Update main with the latest v0.20.0 (36c42c8df)