Run A RayCluster
This page shows how to leverage Kueue’s scheduling and resource management capabilities when running RayCluster.
This guide is for batch users that have a basic understanding of Kueue. For more information, see Kueue’s overview.
Before you begin
Make sure you are using Kueue v0.6.0 version or newer and KubeRay v1.1.0 or newer.
Check Administer cluster quotas for details on the initial Kueue setup.
See KubeRay Installation for installation and configuration details of KubeRay.
Note
In order to use RayCluster, prior to v0.8.1, you need to restart Kueue after the installation. You can do it by running:kubectl delete pods -l control-plane=controller-manager -n kueue-system.RayCluster definition
When running RayClusters on Kueue, take into consideration the following aspects:
a. Queue selection
The target local queue should be specified in the metadata.labels section of the RayCluster configuration.
metadata:
labels:
kueue.x-k8s.io/queue-name: user-queue
b. Configure the resource needs
The resource needs of the workload can be configured in the spec.
spec:
headGroupSpec:
template:
spec:
containers:
- resources:
requests:
cpu: "1"
workerGroupSpecs:
- template:
spec:
containers:
- resources:
requests:
cpu: "1"
Note that a RayCluster will hold resource quotas while it exists. For optimal resource management, you should delete a RayCluster that is no longer in use.
c. Suspend control
Kueue controls the spec.suspend field of the RayCluster. When a RayCluster is admitted by Kueue, Kueue will unsuspend it by setting spec.suspend to false, regardless of its previous value.
d. Limitations
Limited Worker Groups: Because a Kueue workload can have a maximum of 18 PodSets, the maximum number of
spec.workerGroupSpecsis 17In-Tree Autoscaling Constraints: Autoscaling is only supported for elastic RayCluster objects. To enable in-tree autoscaling:
Activate the
ElasticJobsViaWorkloadSlicesfeature gate.Annotate the RayCluster object with:
metadata: annotations: kueue.x-k8s.io/elastic-job: "true"Enable the Ray autoscaler of your RayCluster object by setting:
spec: enableInTreeAutoscaling: true
Example RayCluster
The RayCluster looks like the following:
apiVersion: ray.io/v1
kind: RayCluster
metadata:
name: raycluster-complete
labels:
kueue.x-k8s.io/queue-name: user-queue
spec:
rayVersion: '2.55.1'
headGroupSpec:
serviceType: ClusterIP
rayStartParams:
dashboard-host: '0.0.0.0'
template:
spec:
containers:
- name: ray-head
image: rayproject/ray:2.55.1
resources:
limits:
cpu: "1"
memory: "5Gi"
requests:
cpu: "1"
memory: "5Gi"
ports:
- containerPort: 6379
name: gcs
- containerPort: 8265
name: dashboard
- containerPort: 10001
name: client
lifecycle:
preStop:
exec:
command: ["/bin/sh","-c","ray stop"]
volumeMounts:
- mountPath: /tmp/ray
name: ray-logs
volumes:
- name: ray-logs
emptyDir: {}
workerGroupSpecs:
- replicas: 1
minReplicas: 1
maxReplicas: 10
groupName: small-group
rayStartParams: {}
template:
spec:
containers:
- name: ray-worker
image: rayproject/ray:2.55.1
resources:
limits:
cpu: "1"
memory: "1Gi"
requests:
cpu: "1"
memory: "1Gi"
lifecycle:
preStop:
exec:
command: ["/bin/sh","-c","ray stop"]
volumeMounts:
- mountPath: /tmp/ray
name: ray-logs
volumes:
- name: ray-logs
emptyDir: {}
You can submit a Ray Job using the CLI or log into the Ray Head and execute a job following this example with kind cluster.
Feedback
Was this page helpful?
Glad to hear it! Please tell us how we can improve.
Sorry to hear that. Please tell us how we can improve.