This document describes how to configure and deploy ephemeral RayJobs on Google Kubernetes Engine (GKE) using Kueue concurrent admission to place and migrate the job across compute options.
Running an ephemeral RayJob means that you submit a job definition to Kubernetes, where the operator automatically provisions a dedicated Ray cluster, runs your workload to completion, retrieves the results and status, and immediately deletes the cluster to release resources.
Concurrent admission is a Kueue feature that allows ephemeral RayJobs to start immediately on available resources and fall back across different compute options—such as reservations, Dynamic Workload Scheduler Calendar, Flex-start VMs, on-demand, and Spot VMs—within a single job request. Rather than waiting idly for a specific preferred compute option to become available, a RayJob can start immediately using any available compute option, reducing time-to-start while maximizing overall fleet utilization across your provisioned capacity.
How concurrent admission works
Concurrent admission is an alpha Kueue feature (available in Kueue v0.18 or later) that changes how jobs are placed across compute options:
- Concurrent evaluation: Rather than evaluating compute options serially, Kueue attempts admission across multiple acceptable compute options (also called ResourceFlavors) concurrently. By generating parallel workload variants for each compute option, Kueue schedules and runs the entire job on the first compute option that has available capacity.
- Optional migration: If enabled
(
concurrentAdmissionPolicy.migration.mode: TryPreferredFlavors), Kueue continues attempting admission across more preferred compute options concurrently. If a more preferred compute option becomes available later, Kueue migrates the running job to it. If migration is disabled or omitted, the job remains on its initial compute option to avoid restart disruption. - Workload structure: Kueue implements concurrent admission by marking the original workload as a parent and creating one variant workload per ResourceFlavor. Each variant attempts admission for a single ResourceFlavor independently. When one is admitted, the parent is admitted, and the remaining variants are deactivated.
Before you begin
Before you start, make sure that you have performed the following tasks:
- Enable the Google Kubernetes Engine API. Enable Google Kubernetes Engine API
- To use the Google Cloud CLI for this task,
install and then
initialize the
gcloud CLI. If you previously installed the gcloud CLI, get the latest
version by running the
gcloud components updatecommand. Earlier gcloud CLI versions might not support running the commands in this document.
Make sure you have the following command-line tools installed:
- The Google Cloud CLI (
gcloud) kubectl
Requirements and limitations
- Kueue version: Concurrent admission requires Kueue v0.18 or later.
- Single-compute-option constraint: Concurrent admission requires all PodSets in a workload—including the Ray head and worker group—to be associated with the same single ResourceFlavor. Decoupled compute option assignment (such as scheduling the Ray head node on a standard CPU compute option while placing worker nodes on GPU or TPU compute options) is not supported under concurrent admission policies.
Step 1: Create and connect to the GKE cluster
In this step, you define your environment variables, create a Standard cluster with the Ray operator add-on enabled, create a Compute Engine GPU reservation, add the GPU node pools, and connect to your cluster.
Set environment variables for your project, zone, and cluster name:
export PROJECT_ID=PROJECT_ID export ZONE=ZONE export CLUSTER_NAME=gpu-clusterReplace
PROJECT_IDwith your Google Cloud project ID, andZONEwith your target compute zone (for example,us-central1-a).Create a Standard GKE cluster with the Ray operator add-on enabled:
gcloud container clusters create ${CLUSTER_NAME} \ --project=${PROJECT_ID} \ --zone=${ZONE} \ --machine-type=n2-standard-4 \ --num-nodes=2 \ --addons=RayOperatorCreate a Compute Engine GPU reservation for the GPU node pool to consume. Specify the
--require-specific-reservationflag to prevent non-GKE VMs from consuming the capacity:gcloud compute reservations create my-gpu-reservation \ --project=${PROJECT_ID} \ --zone=${ZONE} \ --vm-count=3 \ --machine-type=n1-standard-4 \ --accelerator=type=nvidia-tesla-t4,count=1 \ --require-specific-reservationCreate the GPU node pools. In this example, you create a GPU pool of reservations and a GPU pool of Flex-start VMs that use NVIDIA T4 GPUs:
Create the reservation GPU node pool:
gcloud container node-pools create gpu-pool-reservation \ --cluster=${CLUSTER_NAME} \ --project=${PROJECT_ID} \ --zone=${ZONE} \ --machine-type=n1-standard-4 \ --accelerator=type=nvidia-tesla-t4,count=1 \ --reservation-affinity=specific \ --reservation=my-gpu-reservation \ --num-nodes=0 \ --enable-autoscaling --min-nodes=0 --max-nodes=3Create the GPU node pool of Flex-start VMs:
gcloud container node-pools create gpu-pool-dws-flex \ --cluster=${CLUSTER_NAME} \ --project=${PROJECT_ID} \ --zone=${ZONE} \ --machine-type=n1-standard-4 \ --accelerator=type=nvidia-tesla-t4,count=1 \ --num-nodes=0 \ --enable-autoscaling --min-nodes=0 --max-nodes=3 \ --flex-start \ --location-policy=ANY \ --reservation-affinity=none
Retrieve cluster credentials to configure
kubectl:gcloud container clusters get-credentials ${CLUSTER_NAME} \ --project=${PROJECT_ID} \ --zone=${ZONE}
Step 2: Install Kueue
Install the Kueue manifests:
kubectl apply --server-side -f https://github.com/kubernetes-sigs/kueue/releases/download/v0.18.0/manifests.yamlEnable the
ConcurrentAdmissionfeature gate by editing thekueue-manager-configConfigMap in thekueue-systemnamespace:kubectl edit configmap kueue-manager-config -n kueue-systemAdd the
featureGatesentry to thecontroller_manager_config.yamldata:apiVersion: v1 kind: ConfigMap metadata: name: kueue-manager-config namespace: kueue-system data: controller_manager_config.yaml: | apiVersion: config.kueue.x-k8s.io/v1beta2 kind: Configuration featureGates: ConcurrentAdmission: true # ... keep the rest of the existing configuration ...Restart the Kueue controller manager deployment to apply the configuration:
kubectl rollout restart deployment kueue-controller-manager -n kueue-system
Step 3: Configure Kueue
Create the ResourceFlavor resources (representing capacity from reservations and Flex-start VMs), a ClusterQueue with a concurrent admission policy, and a LocalQueue.
Important considerations for this configuration include the following:
- The ClusterQueue must use the
queueingStrategy: BestEffortFIFOkey-value pair. TheStrictFIFOvalue for this field is not supported with concurrent admission. - The ClusterQueue must have exactly one
resourceGroupcontaining at most 16 compute options. - Compute options are listed in order of preference: the
reservation-flavorfield is listed first, followed by thedws-flex-flavorfield. - The
spec.concurrentAdmissionPolicyfield is immutable after the ClusterQueue is created. To modify the policy, delete and re-create the ClusterQueue.
Create a file named
kueue-setup.yamlwith the following:Apply the configuration to your cluster:
kubectl apply -f kueue-setup.yaml
Step 4: Define the ephemeral RayJob
With concurrent admission, each admitted variant associates the entire workload with one flavor, and Kueue injects that flavor's node selector onto the head and worker Pods. You define a single worker group (with three worker replicas) and let Kueue decide and migrate which GPU compute option runs it.
Important settings in this definition include the following:
- No compute option node selectors: Do not associate the worker group with a specific node pool. Kueue automatically adds the node selector based on the admitted compute option.
- Head GPU toleration: Because GKE taints GPU nodes with the
nvidia.com/gpu=present:NoScheduletaint, the head Pod requires a matching toleration to colocate on a GPU node with a worker. - Ephemeral behavior: The
shutdownAfterJobFinishes: truekey-value pair instructs KubeRay to delete the Ray cluster when the job finishes. - Disabled autoscaling: Use fixed replicas (
minReplicas == maxReplicas). Autoscaling is not recommended for ephemeral jobs.
Create a file named
ephemeral-rayjob.yamlwith the following:Submit the RayJob to the cluster:
kubectl apply -f ephemeral-rayjob.yaml
Step 5: Monitor and verify concurrent admission
Inspect the parent and variant workload objects:
kubectl get workloadsTo list only the parent workload, run the following command:
kubectl get workloads -l kueue.x-k8s.io/concurrent-admission-parent=trueWatch Pod creation and placement:
kubectl get pods -w -o wideThe head and worker Pods land on nodes matching the admitted compute option. The head colocates on a GPU node with one of the workers, while the short-lived submitter Pod runs on the default CPU pool.
Check the status of the RayJob:
kubectl get rayjobsView the output logs from the running job:
kubectl logs -l job-name=ephemeral-gpu-jobThe output is similar to the following:
Cluster resources: {'node:10.52.3.6': 1.0, 'node:__internal_head__': 1.0, 'memory': 27917287424.0, 'object_store_memory': 8194336357.0, 'CPU': 7.0, 'GPU': 3.0, 'node:10.52.3.7': 1.0, 'node:10.52.4.6': 1.0, 'node:10.52.2.6': 1.0} (gpu_task pid=277, ip=10.52.4.6) Task 0 running on node with IP: 10.52.4.6 (gpu_task pid=275, ip=10.52.2.6) Task 2 running on node with IP: 10.52.2.6 (gpu_task pid=276, ip=10.52.3.7) Task 1 running on node with IP: 10.52.3.7 Tasks successfully executed on nodes: {'10.52.2.6', '10.52.4.6', '10.52.3.7'} Job 'ephemeral-gpu-job' succeededAfter the job succeeds, verify that KubeRay automatically removed the Ray cluster Pods:
kubectl get pods
Step 6: (Optional) Start on Flex-start VMs and migrate to a reservation
This scenario demonstrates how a job starts on fallback capacity (Flex-start VMs) and subsequently upgrades to a reservation when quota becomes available.
- Simulate initial condition (reservation occupied): If the reservation quota is
fully occupied by other workloads when the RayJob is submitted, the
reservation-flavorvariant cannot be admitted. Job starts on fallback compute option (T1): Kueue admits
dws-flex-flavor, and the RayCluster is created ongpu-pool-dws-flex:kubectl get workloadsThe output shows
dws-flex-flavoradmitted asTrue:NAME ADMITTED rayjob-ephemeral-gpu-job-[hash] True rayjob-ephemeral-gpu-job-variant-dws-flex-flavor-[hash] True rayjob-ephemeral-gpu-job-variant-reservation-flavor-[hash]Reservation capacity frees up (T2): When capacity on the reservation becomes available, Kueue admits the
reservation-flavorvariant:kubectl get workloads -wNAME ADMITTED rayjob-ephemeral-gpu-job-[hash] True rayjob-ephemeral-gpu-job-variant-reservation-flavor-[hash] True rayjob-ephemeral-gpu-job-variant-dws-flex-flavor-[hash] FalseJob migrates to reservation (T3): Because the
reservation-flavorcompute option has higher priority, Kueue migrates the parent workload. Migration restarts the RayCluster on the new compute option, and the Pods are re-created on thegpu-pool-reservationcompute option:kubectl get pods -o wide -w
Clean up
To avoid incurring charges to your Google Cloud account for the resources used in this guide, delete the cluster and reservation:
Delete the GKE cluster:
gcloud container clusters delete ${CLUSTER_NAME} \ --project=${PROJECT_ID} \ --zone=${ZONE} \ --quietDelete the Compute Engine reservation:
gcloud compute reservations delete my-gpu-reservation \ --project=${PROJECT_ID} \ --zone=${ZONE} \ --quiet
What's next
- Learn more about Ray on GKE.
- Read about Kueue concurrent admission.
- Explore GPU sharing and provisioning strategies on GKE.