Run A RayService
This page shows how to leverage Kueue’s scheduling and resource management capabilities when running RayService.
Kueue manages the RayService as a top-level job, similar to how Kueue manages the RayJob.
This guide is for serving users that have a basic understanding of Kueue. For more information, see Kueue’s overview.
Before you begin
Make sure you are using Kueue v0.17.0 or newer and KubeRay v1.3.0 or newer.
Check Administer cluster quotas for details on the initial Kueue setup.
See KubeRay Installation for installation and configuration details of KubeRay.
RayService definition
When running RayService on Kueue, take into consideration the following aspects:
a. Queue selection
The target local queue should be specified in the metadata.labels section of the RayService configuration, and this label will be propagated to its RayCluster.
metadata:
labels:
kueue.x-k8s.io/queue-name: user-queue
b. Configure the resource needs
The resource needs of the workload can be configured in the spec.rayClusterConfig.
spec:
rayClusterConfig:
headGroupSpec:
template:
spec:
containers:
- resources:
requests:
cpu: "1"
workerGroupSpecs:
- template:
spec:
containers:
- resources:
requests:
cpu: "1"
c. Suspend control
Kueue controls the spec.rayClusterConfig.suspend field of the RayService. When a RayService is admitted by Kueue, Kueue will unsuspend it by setting spec.rayClusterConfig.suspend to false, regardless of its previous value.
d. Limitations
- Limited Worker Groups: Because a Kueue workload can have a maximum of 18 PodSets, the maximum number of
spec.rayClusterConfig.workerGroupSpecsis 17.
Autoscaling (a.k.a InTreeAutoscaling)
Ray Autoscaling can automatically add or remove Ray worker pods based on resource demand.
This feature is supported for RayService starting in Kueue v0.19.0.
Autoscaling is only supported for elastic RayService objects.
How to enable autoscaling in RayService
Enable the feature gate for Elastic Workloads (Workload Slices):
ElasticJobsViaWorkloadSlices: trueAdd the workload slicing annotation to the RayService object:
metadata: annotations: kueue.x-k8s.io/elastic-job: "true"Enable the Ray autoscaler in the RayService object:
spec: rayClusterConfig: enableInTreeAutoscaling: true
Rolling upgrades limitation
Kueue’s Workload Slices feature currently manages quota for a single active cluster. Upgrade strategies that provision a secondary surge cluster (spec.upgradeStrategy.type: NewCluster or NewClusterWithIncrementalUpgrade) are not currently supported when workload slicing is enabled because pending cluster pods will remain gated. To use workload slicing with autoscaling, use spec.upgradeStrategy.type: None or apply updates in-place.
Example RayService
The RayService looks like the following:
apiVersion: ray.io/v1
kind: RayService
metadata:
name: test-rayservice
namespace: default
labels:
kueue.x-k8s.io/queue-name: user-queue
spec:
upgradeStrategy:
type: None
serveConfigV2: |
applications:
- name: fruit_app
import_path: fruit.deployment_graph
route_prefix: /fruit
runtime_env:
working_dir: "https://github.com/ray-project/test_dag/archive/78b4a5da38796123d9f9ffff59bab2792a043e95.zip"
deployments:
- name: MangoStand
num_replicas: 2
max_replicas_per_node: 1
user_config:
price: 3
ray_actor_options:
num_cpus: 0.1
- name: OrangeStand
num_replicas: 1
user_config:
price: 2
ray_actor_options:
num_cpus: 0.1
- name: PearStand
num_replicas: 1
user_config:
price: 1
ray_actor_options:
num_cpus: 0.1
- name: FruitMarket
num_replicas: 1
ray_actor_options:
num_cpus: 0.1
- name: math_app
import_path: conditional_dag.serve_dag
route_prefix: /calc
runtime_env:
working_dir: "https://github.com/ray-project/test_dag/archive/78b4a5da38796123d9f9ffff59bab2792a043e95.zip"
deployments:
- name: Adder
num_replicas: 1
user_config:
increment: 3
ray_actor_options:
num_cpus: 0.1
- name: Multiplier
num_replicas: 1
user_config:
factor: 5
ray_actor_options:
num_cpus: 0.1
- name: Router
num_replicas: 1
rayClusterConfig:
rayVersion: '2.55.1'
headGroupSpec:
rayStartParams: {}
template:
spec:
containers:
- name: ray-head
image: rayproject/ray:2.55.1
resources:
limits:
cpu: "2"
memory: "5Gi"
requests:
cpu: "2"
memory: "2Gi"
workerGroupSpecs:
- replicas: 1
minReplicas: 1
maxReplicas: 5
groupName: small-group
rayStartParams: {}
template:
spec:
containers:
- name: ray-worker
image: rayproject/ray:2.55.1
resources:
limits:
cpu: "1"
memory: "2Gi"
requests:
cpu: "500m"
memory: "2Gi"
Note
The example above comes from the KubeRay RayService sample and only has thequeue-name label added and requests updated.Feedback
Was this page helpful?
Glad to hear it! Please tell us how we can improve.
Sorry to hear that. Please tell us how we can improve.