This directory contains the Helm chart for deploying the Snapshot Agent DaemonSet in a Kubernetes cluster.
Public images are published to ghcr.io/llm-d-incubation/llm-d-rl-time-slicing/* by CI: latest on every merge to main; versioned tags via a manual workflow run.
- A Kubernetes cluster with GPU nodes (NVIDIA).
kubectlconfigured to connect to your cluster.helm(v3+) installed.
Important
The Snapshot Agent is hardcoded to be deployed in the timeslice-system namespace. Consequently, the Helm chart creates resources specifically in the timeslice-system namespace.
To deploy the agent independently using the local Helm chart:
-
Install the chart: From the
deploydirectory, install the chart into thetimeslice-systemnamespace (creating it if it doesn't exist):helm install snapshot-agent ./snapshot-agent \ --namespace timeslice-system \ --create-namespace
This will deploy the agent as a
DaemonSetand set up the required RBAC permissions:- Creating a
ServiceAccountfor the agent. - Creating a
ClusterRoleandClusterRoleBindinggranting permissions toget,list, andwatchpods and nodes, andgetonnodes/proxy. - Configuring the agent pods to use this
ServiceAccount.
- Creating a
-
Verify the deployment:
kubectl get pods -n timeslice-system -l app.kubernetes.io/name=snapshot-agent
-
Uninstall the chart:
helm uninstall snapshot-agent --namespace timeslice-system
- A GKE cluster with at least one GPU node pool.
- The NVIDIA GPU device driver must be installed on the nodes (e.g., using the GKE GPU driver installer).
nvidia.driver.hostPath:/home/kubernetes/bin/nvidia(Standard path for GPU drivers on GKE COS).nvidia.devices.hostPath:/dev(Standard path for device access).tolerations: Includesnvidia.com/gputo allow the agent to run on GPU-tainted nodes.
To install the chart on GKE, ensuring it only targets nodes with GPUs:
helm install snapshot-agent ./snapshot-agent \
--namespace timeslice-system \
--create-namespace \
--set-string "nodeSelector.cloud\.google\.com/gke-gpu=true"If your GKE nodes are using Ubuntu instead of COS, you may need to override the driver path:
helm install snapshot-agent ./snapshot-agent \
--namespace timeslice-system \
--create-namespace \
--set-string "nodeSelector.cloud\.google\.com/gke-gpu=true" \
--set nvidia.driver.hostPath=/usr/lib/nvidiaIf you are deploying the snapshot agent to a non-GKE cluster (e.g., EKS, AKS, or bare-metal), you will likely need to adjust the node selector and driver paths because they differ from GKE defaults.
Non-GKE clusters typically use different labels to identify GPU nodes. For example, standard NVIDIA GPU nodes often use nvidia.com/gpu=true or hardware=gpu.
You can override the GKE-default node selector during installation:
helm install snapshot-agent ./snapshot-agent \
--namespace timeslice-system \
--create-namespace \
--set-string "nodeSelector.nvidia\.com/gpu=true"Note: You may need to escape the dots in the label key as shown above (nodeSelector.nvidia\.com/gpu=true).
On non-GKE clusters, the NVIDIA driver libraries might be installed in different locations on the host. Common paths include:
/usr/lib/nvidia/usr/local/nvidia/usr/lib/x86_64-linux-gnu
You can override the host path using:
helm install snapshot-agent ./snapshot-agent \
--namespace timeslice-system \
--create-namespace \
--set nvidia.driver.hostPath=/usr/lib/nvidiaIf your GPU nodes have different taints than the default nvidia.com/gpu=present:NoSchedule, you must override the tolerations. For example, if your nodes are tainted with sku=gpu:NoSchedule:
helm install snapshot-agent ./snapshot-agent \
--namespace timeslice-system \
--create-namespace \
--set tolerations[0].key=sku \
--set tolerations[0].operator=Equal \
--set tolerations[0].value=gpu \
--set tolerations[0].effect=NoScheduleDuring development, you will need to build your own container image containing your changes and push it to a custom registry.
We use the provided Makefile targets to build and push the container image.
- Define your custom registry and version (tag) by setting them as environment variables:
export REGISTRY=your-custom-registry.com/your-project export VERSION=dev-$(git rev-parse --short HEAD)
- Run the following make target from the repository root to build and push the image:
This will build the image and push it to
make snapshot-agent-image-push
your-custom-registry.com/your-project/llm-d-rl-time-slicing/snapshot-agent:dev-<hash>(the repo name comes fromPROJECT_NAME, also overridable).
Once your image is pushed, you can instruct Helm to use it.
helm install snapshot-agent ./snapshot-agent \
--namespace timeslice-system \
--create-namespace \
--set image.repository=your-custom-registry.com/your-project/snapshot-agent \
--set image.tag=devEdit deploy/snapshot-agent/values.yaml directly:
image:
repository: your-custom-registry.com/your-project/snapshot-agent
pullPolicy: IfNotPresent
tag: "dev"And then run:
helm install snapshot-agent ./snapshot-agent \
--namespace timeslice-system \
--create-namespaceThe direct_memory backend parks and resumes a workload's entire GPU
state via GPU-CR. The workload's preloader drains the device memory it
tracked into node-local hugepage-backed files over a pinned DMA path
before the CUDA context is frozen with cuda-checkpoint, so the bulk
bytes bypass that tool's slow pageable copy — park and resume run several
times faster for the same VRAM, and the parked state lives in files that
survive independently of the process's memory.
It is gated behind the directMemory values block and disabled by
default — with directMemory.enabled=false the rendered chart is
identical to a plain CUDA/app-backend deployment.
helm install snapshot-agent ./snapshot-agent \
--namespace timeslice-system \
--create-namespace \
--set directMemory.enabled=trueEnabling it adds:
- The checkpoint directory shared with workloads, mounted at
directMemory.ctlDir(default/mnt/huge-ckpt) fromdirectMemory.hostCtlPath(default/var/tmp/huge-ckpt), withmountPropagation: HostToContainerso the init containers' mounts are visible. The agent'sEXPORT_FILE_PATHand per-operation deadline (DIRECT_MEMORY_OP_TIMEOUT_SEC,opTimeoutSec, default 120 s) are set from this block. - The
DirectMemoryBackendfeature gate, implied bydirectMemory.enabled=true(an explicitfeatureGatesentry overrides the implied value). - A
PriorityClass(priorityClass.*) so the pod — and in particular its hugepage bootstrap — wins node placement over GPU workloads. - Three privileged init containers that
nsenterthe host mount namespace, all idempotent per node boot (rendered whiledirectMemory.hugetlbfs.mountis on, the default):provision-hugepages(hugetlbfs.bootstrap.*) — writesvm.nr_hugepages(pages2Mi, default 12288 = 24 Gi; size it to your workloads' dump buffers plus headroom) and restarts the kubelet once so the node publisheshugepages-2Micapacity for WORKLOAD pods. No nodepool hugepage configuration or special node image is needed.mount-hugetlbfs— mounts hugetlbfs athostCtlPath(pagesize=2M,mode=0777). GPU-CR pins its dump and staging files for DMA, which requires hugepage-backed files; without this mount every dump silently degrades to boot-disk page cache. Turnhugetlbfs.mountoff only if the path already is a hugetlbfs.mount-ctl-tmpfs— mounts a small tmpfs (ctlTmpfsSizeMi, default 64) nested at<hostCtlPath>/ctlfor the control files through which the agent'scr_clientand the workload's preloader coordinate. Both sides discover it through the store mount they already share — no configuration on either side — and keeping control files off hugetlbfs is what lets the agent run with no hugepage request.
The agent requests no hugepages-2Mi at all: dump bytes are written by
the workload's own process, and the agent's control traffic stays on the
tmpfs. That zero request is what lets the DaemonSet schedule on fresh nodes
before hugepage capacity exists and absorb the hugepage bootstrap as an
init container. (directMemory.hugepagesResource exists only for older
GPU-CR builds that keep control files on the store, on nodes whose pool is
already provisioned.)
Misconfigurations fail early:
- An undersized or unallocatable hugepage pool makes
provision-hugepagesexit 1 without the kubelet restart — it surfaces as this pod CrashLooping, never as workload SIGBUS. hugetlbfs.bootstrapcombined with a pagesize other than2M, or withhugepagesResource(a pod that requests hugepages cannot schedule before its own bootstrap publishes capacity), fails at render time.
Workload pods need the GPU-CR preloader (LD_PRELOAD=vGPU-NVIDIA.so),
hostPID, the checkpoint dir hostPath mounted at /mnt/huge-ckpt with
mountPropagation: HostToContainer, and hugepages-2Mi resources sized for
their dump buffers — see the workload example in
guides/snapshot-agent.
There is no cr_client install step: the binary is built from
third_party/gpu-cr into the agent image at /usr/local/bin/cr_client, so
the agent, cr_client, and the preloader source always roll together.
grpc.health.v1.Health/Check with service: "direct-memory" reports
NOT_SERVING if the binary is missing.
Note
Clusters that previously ran a manually deployed agent may already have a
non-Helm timeslice-snapshot-agent PriorityClass; helm install refuses
to adopt it. Delete it once before installing:
kubectl delete priorityclass timeslice-snapshot-agent.