| sidebar-title | Getting Started on Kubernetes |
|---|
AI Agents (Claude, Copilot, Cursor, etc.): For diagnosing failures, see the AI Agent Debugging Guide.
This guide walks you through benchmarking an NVIDIA Dynamo inference server on Kubernetes using AIPerf. By the end, you will have a cluster running, both operators installed, a benchmark executed against a Dynamo deployment, and your results downloaded.
The same workflow applies to any OpenAI-compatible endpoint (vLLM, TRT-LLM, SGLang) -- just change the endpoint URL and model name.
If you already have a Kubernetes cluster with GPU nodes and kubectl configured, skip to Prerequisites.
Any Kubernetes v1.24+ cluster with NVIDIA GPUs works. You need:
kubectlconfigured to talk to the cluster- GPU nodes with the NVIDIA device plugin installed
- Permissions to install CRDs and create namespaces
Verify your connection:
kubectl cluster-info
kubectl get nodes -o custom-columns=NAME:.metadata.name,GPU:.status.allocatable.nvidia\\.com/gpuIf GPUs show up, skip to Prerequisites.
Kind runs Kubernetes inside Docker containers on your local machine. With the NVIDIA container runtime configured as Docker's default, a Kind node can see the host's GPUs.
Host requirements:
- Docker with the NVIDIA container runtime as the default runtime
- NVIDIA drivers installed on the host
kindCLI installed (go install sigs.k8s.io/kind@latestor releases)
If you haven't already configured Docker for GPU passthrough:
# Enable volume-mount-based GPU injection
sudo nvidia-ctk config --in-place \
--set accept-nvidia-visible-devices-as-volume-mounts=true
# Set nvidia as the default Docker runtime
sudo nvidia-ctk runtime configure --runtime=docker --set-as-default
# Restart Docker to pick up changes
sudo systemctl restart dockerVerify:
docker info 2>/dev/null | grep "Default Runtime"
# Should show: Default Runtime: nvidiaCreate the cluster:
kind create cluster --name aiperfInstall the NVIDIA device plugin so GPUs become an allocatable resource, and JobSet, which the AIPerf operator uses to run benchmark pods:
kubectl --context kind-aiperf apply -f \
https://raw.githubusercontent.com/NVIDIA/k8s-device-plugin/master/deployments/static/nvidia-device-plugin.yml
kubectl --context kind-aiperf apply --server-side -f \
https://github.com/kubernetes-sigs/jobset/releases/latest/download/manifests.yamlBuild the AIPerf image locally and load it into the Kind node so pods can pull it without a registry:
docker build -t aiperf:local .
kind load docker-image aiperf:local --name aiperfVerify GPUs are allocatable:
kubectl --context kind-aiperf get nodes \
-o jsonpath='{.items[0].status.allocatable.nvidia\.com/gpu}'
# Should show: 1 (or more)To delete the cluster when you are done:
kind delete cluster --name aiperfAt this point you should have:
- A Kubernetes cluster with
kubectlconfigured - GPU nodes with the NVIDIA device plugin
- JobSet installed
- Helm v3 installed locally
- AIPerf installed locally (
uv tool install aiperf, oruv syncfrom a source checkout) - Access to NGC container registry (
nvcr.io/nvidia/ai-dynamo)
Run the preflight checker to verify:
aiperf kube preflightYou need two operators: the Dynamo operator (manages inference servers) and the AIPerf operator (manages benchmarks).
# Install Dynamo platform (CRDs are bundled into the platform chart in v1.x)
helm install dynamo-platform \
oci://nvcr.io/nvidia/ai-dynamo/dynamo-platform \
--version 1.1.0 \
--namespace dynamo-system \
--create-namespace \
--set dynamo-operator.webhook.enabled=false \
--set grove.enabled=false \
--set kai-scheduler.enabled=falseVerify the Dynamo operator is running:
kubectl get pods -n dynamo-systemhelm install aiperf-operator deploy/helm/aiperf-operator \
--namespace aiperf-system \
--create-namespaceFor Kind clusters using a locally built image, override the image:
helm install aiperf-operator deploy/helm/aiperf-operator \
--namespace aiperf-system \
--create-namespace \
--set image.repository=aiperf \
--set image.tag=local \
--set image.pullPolicy=NeverVerify it is running:
kubectl get pods -n aiperf-systemYou should see 2/2 containers ready by default (operator + results-server). With the optional Plotly dashboard enabled (dashboard.enabled=true), the count becomes 3/3 (operator + results-server + dashboard).
Create a DynamoGraphDeployment. This example deploys Qwen3-0.6B in aggregated mode using the vLLM backend:
# dynamo-server.yaml
apiVersion: nvidia.com/v1alpha1
kind: DynamoGraphDeployment
metadata:
name: dynamo-agg
namespace: dynamo-server
spec:
services:
Frontend:
dynamoNamespace: dynamo-agg
componentType: frontend
replicas: 1
extraPodSpec:
mainContainer:
image: nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.1.0
env:
- name: POD_UID
valueFrom:
fieldRef:
fieldPath: metadata.uid
VllmWorker:
dynamoNamespace: dynamo-agg
componentType: worker
replicas: 1
resources:
limits:
gpu: "1"
extraPodSpec:
runtimeClassName: nvidia
mainContainer:
image: nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.1.0
workingDir: /workspace/examples/backends/vllm
command: ["python3", "-m", "dynamo.vllm"]
args:
- "--model"
- "Qwen/Qwen3-0.6B"
env:
- name: POD_UID
valueFrom:
fieldRef:
fieldPath: metadata.uid
- name: DYN_SYSTEM_PORT
value: "9090"Apply it:
kubectl create namespace dynamo-server
kubectl apply -f dynamo-server.yamlWait for the server to be ready (model loading can take a few minutes):
# Watch pods until all are Running
kubectl get pods -n dynamo-server -wVerify the endpoint is healthy:
kubectl run curl-test --rm -it --restart=Never --image=curlimages/curl -- \
curl -s http://dynamo-agg-frontend.dynamo-server.svc:8000/v1/modelsNow benchmark the Dynamo server. The Dynamo endpoint URL follows the pattern http://{deployment-name}-frontend.{namespace}.svc:8000/v1:
aiperf kube profile \
--model Qwen/Qwen3-0.6B \
--url http://dynamo-agg-frontend.dynamo-server.svc:8000/v1 \
--image nvcr.io/nvidia/aiperf:latest \
--total-workers 10 \
--request-count 500 \
--concurrency 50 \
--streamingOn a Kind cluster with a locally built image, point --image at the loaded tag and disable pulling. Substitute the Service URL of whatever OpenAI-compatible endpoint you are testing against:
aiperf kube profile \
--model Qwen/Qwen3-0.6B \
--url http://my-endpoint.default.svc:8000 \
--image aiperf:local \
--image-pull-policy Never \
--total-workers 10 \
--request-count 200 \
--concurrency 5 \
--streamingWhat happens:
- AIPerf builds an
AIPerfJobcustom resource from your flags - Submits it to the cluster, where the AIPerf operator picks it up
- The operator creates RBAC, a ConfigMap, and a JobSet with a controller pod and worker pods
- Workers send requests to the Dynamo frontend
- AIPerf polls the
AIPerfJobstatus and streams progress to your terminal
You will see the job phase, worker readiness (ready/total), and each status condition as the operator reports it. In --no-operator mode AIPerf attaches to the controller pod and tails its output instead.
Press Ctrl+C to detach. The benchmark continues running in the cluster. To cancel it, run aiperf kube cancel or patch the CR directly:
kubectl patch aiperfjob <name> -n my-benchmarks --type=merge -p '{"spec":{"cancel":true}}'For repeatable benchmarks, use an AIPerfJob YAML file. Generate a starter template:
aiperf kube init --output benchmark.yamlEdit it for your Dynamo deployment:
apiVersion: aiperf.nvidia.com/v1alpha1
kind: AIPerfJob
metadata:
name: dynamo-benchmark
spec:
benchmark:
models:
- "Qwen/Qwen3-0.6B"
endpoint:
urls:
- "http://dynamo-agg-frontend.dynamo-server.svc:8000/v1"
streaming: true
datasets:
- name: main
type: synthetic
entries: 1000
prompts:
isl:
mean: 512
stddev: 0
osl:
mean: 128
stddev: 0
phases:
- name: warmup
kind: warmup
type: concurrency
concurrency: 10
requests: 20
- name: profiling
kind: profiling
type: concurrency
concurrency: 50
requests: 500Validate the config before deploying:
aiperf kube validate benchmark.yamlRun it:
aiperf kube profile --config benchmark.yaml --image nvcr.io/nvidia/aiperf:latestOr apply it directly with kubectl (the operator picks it up automatically):
kubectl apply -f benchmark.yamlDynamo's disaggregated mode separates prefill and decode into different pods for better GPU utilization. To benchmark a disaggregated deployment:
- Deploy Dynamo in disaggregated mode (separate prefill and decode workers with KV cache transfer):
# dynamo-disagg.yaml
apiVersion: nvidia.com/v1alpha1
kind: DynamoGraphDeployment
metadata:
name: dynamo-disagg
namespace: dynamo-server
spec:
services:
Frontend:
dynamoNamespace: dynamo-disagg
componentType: frontend
replicas: 1
extraPodSpec:
mainContainer:
image: nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.1.0
env:
- name: POD_UID
valueFrom:
fieldRef:
fieldPath: metadata.uid
- name: DYN_ROUTER_MODE
value: "kv"
VllmPrefillWorker:
dynamoNamespace: dynamo-disagg
componentType: worker
subComponentType: prefill
replicas: 1
resources:
limits:
gpu: "1"
extraPodSpec:
runtimeClassName: nvidia
mainContainer:
image: nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.1.0
workingDir: /workspace/examples/backends/vllm
command: ["python3", "-m", "dynamo.vllm"]
args:
- "--model"
- "Qwen/Qwen3-0.6B"
- "--is-prefill-worker"
- "--connector"
- "kvbm"
env:
- name: POD_UID
valueFrom:
fieldRef:
fieldPath: metadata.uid
- name: DYN_KVBM_CPU_CACHE_GB
value: "1"
VllmDecodeWorker:
dynamoNamespace: dynamo-disagg
componentType: worker
subComponentType: decode
replicas: 1
resources:
limits:
gpu: "1"
extraPodSpec:
runtimeClassName: nvidia
mainContainer:
image: nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.1.0
workingDir: /workspace/examples/backends/vllm
command: ["python3", "-m", "dynamo.vllm"]
args:
- "--model"
- "Qwen/Qwen3-0.6B"
- "--is-decode-worker"
env:
- name: POD_UID
valueFrom:
fieldRef:
fieldPath: metadata.uid
- name: DYN_SYSTEM_PORT
value: "9090"- Benchmark it -- the endpoint URL changes to match the deployment name:
aiperf kube profile \
--model Qwen/Qwen3-0.6B \
--url http://dynamo-disagg-frontend.dynamo-server.svc:8000/v1 \
--image nvcr.io/nvidia/aiperf:latest \
--total-workers 20 \
--request-count 1000 \
--concurrency 100 \
--streamingAfter the benchmark completes, retrieve your results:
aiperf kube resultsThis downloads the full results package including:
profile_export_aiperf.json-- Summary metrics (throughput, latency percentiles, TTFT, ITL)inputs.json-- Dataset that was usedserver_metrics_export.json-- Dynamo server metrics (frontend throughput, KV cache stats, component latencies); empty when discovery found no scrapable endpoints
To retrieve results from a specific job:
aiperf kube results dynamo-benchmarkResults are stored on the operator's persistent volume by default, so aiperf kube results works even after benchmark pods are deleted. To retrieve directly from the benchmark pods instead (downloads every artifact through the controller's results API, so the controller pod must still be running):
aiperf kube results dynamo-benchmark --from-podsAdding --summary-only narrows the download to the summary files and falls back to kubectl cp when the controller API is unreachable.
When benchmarking Dynamo, AIPerf discovers and collects Prometheus metrics from pods with the nvidia.com/metrics-enabled=true label (or a recognizable inference-server image). By default discovery only searches the benchmark job's own namespace — the namespace the chart-provisioned benchmark RBAC can list pods in. If Dynamo runs in a different namespace (e.g. dynamo-server), set server_metrics.discovery.namespace: dynamo-server in the benchmark config and grant pod-read access there by adding that namespace to the chart's serverMetricsDiscoveryNamespaces value — a plain string entry binds the benchmark namespaces' default ServiceAccount, so benchmark pods running under a custom podTemplate.serviceAccountName need the {namespace, serviceAccounts} entry form (or a manual RoleBinding); see the server-metrics guide for the full RBAC prerequisite table and how to tune or disable discovery. Discovered metrics include:
- Frontend metrics --
dynamo_frontend_requests,dynamo_frontend_time_to_first_token_seconds,dynamo_frontend_inter_token_latency_seconds,dynamo_frontend_output_tokens - Component metrics -- Per-worker
dynamo_component_requests,dynamo_component_kvstats_gpu_cache_usage_percent,dynamo_component_kvstats_gpu_prefix_cache_hit_rate
These are exported alongside the standard AIPerf metrics in the results package.
See all benchmark jobs across namespaces:
aiperf kube listNAME NAMESPACE OWNER PHASE WORKERS PROGRESS THROUGHPUT LATENCY AGE
dynamo-benchmark my-benchmarks - Completed 10/10 100% 142.3 rps 187.0 ms 5m
disagg-test my-benchmarks - Running 10/10 42% 98.1 rps 201.4 ms 2m
WORKERS is ready/total and LATENCY is the p99 request latency. OWNER is
the scoped operator holding that namespace's claim, - when the cluster-wide
operator reconciles it, or ? when the claim could not be read. Use --wide to add model, endpoint, and error columns.
Filter by status:
aiperf kube list --running
aiperf kube list --completed
aiperf kube list --failedWatch jobs with live refresh:
aiperf kube list --watchThe AIPerf operator includes a built-in web dashboard for monitoring benchmarks and analyzing results.
Access it by port-forwarding to the operator:
kubectl port-forward -n aiperf-system deploy/aiperf-operator 8081:8081Then open http://localhost:8081 in your browser.
The dashboard provides:
- Dashboard -- Overview with KPI cards, active jobs, and throughput trends
- Jobs -- Sortable table of all benchmark jobs with phase filters
- Job Detail -- Live metrics, charts, phase progress, and pod status for a single job
- Leaderboard -- Rank benchmark runs by any metric
- Compare -- Side-by-side comparison of multiple jobs
- History -- Time-series charts showing metrics across runs
- Sweeps -- AIPerfSweep listings and per-sweep variation drill-down
Use Ctrl+K to quickly search and navigate between jobs and pages.
| Mode | Description | Endpoint URL Pattern |
|---|---|---|
agg |
Aggregated -- single workers handle prefill + decode | http://dynamo-agg-frontend.dynamo-server.svc:8000/v1 |
agg-router |
Aggregated with KV-aware routing | http://dynamo-agg-router-frontend.dynamo-server.svc:8000/v1 |
disagg |
Disaggregated -- separate prefill and decode workers | http://dynamo-disagg-frontend.dynamo-server.svc:8000/v1 |
Dynamo supports three inference backends. Change the worker image and command:
| Backend | Image | Worker Command |
|---|---|---|
| vLLM | nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.1.0 |
python3 -m dynamo.vllm |
| TRT-LLM | nvcr.io/nvidia/ai-dynamo/trtllm-runtime:1.1.0 |
python3 -m dynamo.trtllm |
| SGLang | nvcr.io/nvidia/ai-dynamo/sglang-runtime:1.1.0 |
python3 -m dynamo.sglang |
If you cannot install the AIPerf operator (e.g., limited cluster permissions), AIPerf can deploy benchmarks directly using raw Kubernetes manifests. Use --no-operator:
aiperf kube profile \
--model Qwen/Qwen3-0.6B \
--url http://dynamo-agg-frontend.dynamo-server.svc:8000/v1 \
--image nvcr.io/nvidia/aiperf:latest \
--no-operatorThis creates RBAC, ConfigMap, and JobSet directly in an existing namespace. You lose operator features (automated monitoring, results storage, conditions) but the benchmark itself works the same way.
To generate the manifests without applying them (useful for GitOps):
aiperf kube generate --no-operator \
--model Qwen/Qwen3-0.6B \
--url http://dynamo-agg-frontend.dynamo-server.svc:8000/v1 \
--image nvcr.io/nvidia/aiperf:latest \
> manifests.yaml| Task | Command |
|---|---|
| Check cluster readiness | aiperf kube preflight |
| Generate config template | aiperf kube init |
| Validate a config file | aiperf kube validate benchmark.yaml |
| Run a benchmark | aiperf kube profile --config benchmark.yaml --image <img> |
| Run without waiting | aiperf kube profile ... --detach |
| Preview the operator CR without deploying | aiperf kube profile ... --dry-run |
| Preview direct-mode manifests without deploying | aiperf kube profile ... --no-operator --dry-run |
| Attach to a running job | aiperf kube attach |
| Diagnose a stuck or failed job | aiperf kube debug |
| List all jobs | aiperf kube list |
| Get results | aiperf kube results |
| Get logs | aiperf kube logs |
| Cancel a running job | aiperf kube cancel |
- Deploy from a Source Checkout -- Build and push AIPerf, install the operator with Helm, and run on a real cluster
- Kubernetes Configuration Reference -- All CRD fields, deployment options, and config patterns
- Monitoring and Troubleshooting -- Watch, debug, and diagnose benchmark issues
- Production Deployments -- CI/CD, Kueue scheduling, secrets, and GitOps workflows
- YAML Config Reference -- General AIPerf YAML configuration
- Sequence Length Distributions -- ISL/OSL distribution configuration
- Architecture -- AIPerf internal architecture