| sidebar-title | Source Checkout Deployment |
|---|
This guide starts from an AIPerf source checkout and deploys that code to a real Kubernetes cluster. It builds a container image, pushes it to a registry, installs or upgrades the AIPerf operator with Helm, then runs a benchmark against a real inference endpoint.
This guide does not use the mock server path. Use it when you have a real OpenAI-compatible endpoint already running, or when you want to deploy a real Dynamo/vLLM endpoint first.
You need:
- An AIPerf source checkout.
kubectlconfigured for the target cluster.- Helm v3.
- Docker or another container builder that can build and push OCI images.
- Registry credentials for the image repository you will use.
- Permission to install CRDs, create the operator namespace, create benchmark namespaces, and create JobSet workloads.
- JobSet installed on the cluster. If it is not installed, install it before the AIPerf operator.
- A real OpenAI-compatible inference endpoint reachable from benchmark pods, or permission to deploy one.
Check the cluster first:
kubectl cluster-info
kubectl get nodes
kubectl api-resources | grep -i jobsetFor GPU benchmarks, verify GPU resources are allocatable:
kubectl get nodes -o custom-columns=NAME:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpuInstall the local development environment from the checkout:
make first-time-setupUse the local CLI through uv run until the package is installed somewhere else:
uv run aiperf --help
uv run aiperf kube --helpPick a tag that identifies the exact checkout you are deploying:
export AIPERF_TAG="$(git rev-parse --short HEAD)"The same image is used by the operator containers and benchmark JobSet pods unless you explicitly override defaults.image in the Helm chart.
Use this shape for an NGC organization or private NVIDIA registry namespace:
export AIPERF_IMAGE="nvcr.io/<org>/aiperf:${AIPERF_TAG}"
docker build -t "${AIPERF_IMAGE}" .
docker push "${AIPERF_IMAGE}"Use this shape for GHCR:
export AIPERF_IMAGE="ghcr.io/<org>/aiperf:${AIPERF_TAG}"
docker build -t "${AIPERF_IMAGE}" .
docker push "${AIPERF_IMAGE}"If your cluster nodes use a different architecture than the build host, build for the cluster platform:
docker buildx build \
--platform linux/amd64 \
-t "${AIPERF_IMAGE}" \
--push \
.Skip this section if every node can pull the image without a secret.
Create the operator namespace first:
kubectl create namespace aiperf-system --dry-run=client -o yaml | kubectl apply -f -For NGC-style registries:
kubectl create secret docker-registry aiperf-registry \
--namespace aiperf-system \
--docker-server=nvcr.io \
--docker-username='$oauthtoken' \
--docker-password="${NGC_API_KEY}"For GHCR:
kubectl create secret docker-registry aiperf-registry \
--namespace aiperf-system \
--docker-server=ghcr.io \
--docker-username="${GITHUB_USER}" \
--docker-password="${GITHUB_TOKEN}"The Helm install below references this secret for the operator. Benchmark jobs also need pull access in the namespace you run them in. Kubernetes secrets are namespace-scoped, so create the same pull secret there:
kubectl create namespace <your-namespace> --dry-run=client -o yaml | kubectl apply -f -
kubectl create secret docker-registry aiperf-registry \
--namespace <your-namespace> \
--docker-server=<registry-host> \
--docker-username=<username> \
--docker-password=<token>Split the image into repository and tag for Helm:
export AIPERF_IMAGE_REPOSITORY="${AIPERF_IMAGE%:*}"
export AIPERF_IMAGE_TAG="${AIPERF_IMAGE##*:}"Install the operator:
helm upgrade --install aiperf-operator deploy/helm/aiperf-operator \
--namespace aiperf-system \
--create-namespace \
--set image.repository="${AIPERF_IMAGE_REPOSITORY}" \
--set image.tag="${AIPERF_IMAGE_TAG}" \
--set image.pullPolicy=IfNotPresentIf you created aiperf-registry, include it in the Helm release:
helm upgrade --install aiperf-operator deploy/helm/aiperf-operator \
--namespace aiperf-system \
--create-namespace \
--set image.repository="${AIPERF_IMAGE_REPOSITORY}" \
--set image.tag="${AIPERF_IMAGE_TAG}" \
--set image.pullPolicy=IfNotPresent \
--set 'imagePullSecrets[0].name=aiperf-registry'The chart default for defaults.image is empty, which means AIPerfJob pods use <image.repository>:<image.tag>, falling back to the chart's appVersion when image.tag is also empty. Set defaults.image only when benchmark pods should run a different image from the operator.
Wait for the operator and results server:
kubectl rollout status deploy/aiperf-operator -n aiperf-system --timeout=180s
kubectl get pods -n aiperf-systemRun Helm tests if the cluster can pull the chart's test image:
helm test aiperf-operator -n aiperf-systemUse this path when your inference server is already deployed in the cluster or reachable from the cluster network.
Set the endpoint and model:
export MODEL="Qwen/Qwen3-0.6B"
export ENDPOINT_URL="http://vllm.default.svc.cluster.local:8000/v1"Run cluster-side preflight checks:
uv run aiperf kube preflight \
--image "${AIPERF_IMAGE}" \
--endpoint-url "${ENDPOINT_URL}" \
--workers 8Submit a benchmark:
uv run aiperf kube profile \
--model "${MODEL}" \
--url "${ENDPOINT_URL}" \
--image "${AIPERF_IMAGE}" \
--total-workers 8 \
--request-count 1000 \
--concurrency 50 \
--streamingFor private benchmark images, pass the pull secret name when submitting jobs:
uv run aiperf kube profile \
--model "${MODEL}" \
--url "${ENDPOINT_URL}" \
--image "${AIPERF_IMAGE}" \
--image-pull-secrets aiperf-registry \
--total-workers 8 \
--request-count 1000 \
--concurrency 50 \
--streamingFor repeatable runs, put the benchmark configuration in YAML and pass --config benchmark.yaml; see End-to-End Workflow for the init -> validate -> preflight -> profile sequence.
Skip this section if you already have a real endpoint.
Install the Dynamo platform chart if your cluster does not already have it:
helm upgrade --install dynamo-platform \
oci://nvcr.io/nvidia/ai-dynamo/dynamo-platform \
--version 1.1.0 \
--namespace dynamo-system \
--create-namespace \
--set dynamo-operator.webhook.enabled=false \
--set grove.enabled=false \
--set kai-scheduler.enabled=falseDeploy an aggregated Dynamo vLLM server by applying a DynamoGraphDeployment manifest such as the aggregated vLLM example in Getting Started on Kubernetes:
kubectl create namespace dynamo-server --dry-run=client -o yaml | kubectl apply -f -
kubectl apply -f dynamo-server.yamlWait for the frontend and worker pods to become ready:
kubectl get pods -n dynamo-server -wUse the Dynamo service URL as the benchmark endpoint:
export ENDPOINT_URL="http://dynamo-agg-frontend.dynamo-server.svc:8000/v1"
uv run aiperf kube preflight \
--image "${AIPERF_IMAGE}" \
--endpoint-url "${ENDPOINT_URL}" \
--workers 8
uv run aiperf kube profile \
--model "${MODEL}" \
--url "${ENDPOINT_URL}" \
--image "${AIPERF_IMAGE}" \
--total-workers 8 \
--request-count 1000 \
--concurrency 50 \
--streamingWatch progress:
uv run aiperf kube listReattach to a detached run:
uv run aiperf kube attachDownload results:
uv run aiperf kube results --output ./aiperf-resultsPort-forward the results server and dashboard:
kubectl port-forward -n aiperf-system svc/aiperf-operator 8081:8081Then open http://localhost:8081.
After changing source code, repeat the build and push with a new immutable tag:
export AIPERF_TAG="$(git rev-parse --short HEAD)-$(date +%Y%m%d%H%M%S)"
export AIPERF_IMAGE="nvcr.io/<org>/aiperf:${AIPERF_TAG}"
docker build -t "${AIPERF_IMAGE}" .
docker push "${AIPERF_IMAGE}"
export AIPERF_IMAGE_REPOSITORY="${AIPERF_IMAGE%:*}"
export AIPERF_IMAGE_TAG="${AIPERF_IMAGE##*:}"
helm upgrade aiperf-operator deploy/helm/aiperf-operator \
--namespace aiperf-system \
--reuse-values \
--set image.repository="${AIPERF_IMAGE_REPOSITORY}" \
--set image.tag="${AIPERF_IMAGE_TAG}"Wait for the new operator pod before submitting new jobs:
kubectl rollout status deploy/aiperf-operator -n aiperf-system --timeout=180sCheck the failing pod and events:
kubectl describe pod -n aiperf-system -l app.kubernetes.io/name=aiperf-operator
uv run aiperf kube debugCommon fixes:
- Confirm the image was pushed with the exact tag used by Helm or
aiperf kube profile. - Create the pull secret in both
aiperf-systemand the benchmark namespace. - Pass
--image-pull-secrets aiperf-registrytoaiperf kube profilefor private benchmark images. - Use
--set image.pullPolicy=Alwayswhile testing mutable tags; prefer immutable tags for normal use.
The operator image and default benchmark image come from the Helm release. Check the rendered CRD default:
kubectl get crd aiperfjobs.aiperf.nvidia.com \
-o jsonpath='{.spec.versions[?(@.name=="v1alpha1")].schema.openAPIV3Schema.properties.spec.properties.image.default}{"\n"}'If you set defaults.image, it overrides the benchmark image independently of image.repository and image.tag.
Verify the endpoint from inside the cluster:
kubectl run curl-check \
--rm -it \
--restart=Never \
--image=curlimages/curl:latest \
-- curl -sf "${ENDPOINT_URL}/models"If the server is still starting and you intentionally want the benchmark to wait until worker runtime, set skipEndpointCheck: true in the AIPerfJob YAML. Do not use this to hide a wrong service name or namespace.
The Helm chart creates benchmark RBAC in the release namespace and in any namespaces listed in benchmarkRbacNamespaces. If you run jobs in a different namespace, add that namespace to benchmarkRbacNamespaces:
helm upgrade aiperf-operator deploy/helm/aiperf-operator \
--namespace aiperf-system \
--reuse-values \
--set 'benchmarkRbacNamespaces[0]=team-a-benchmarks'- Getting Started on Kubernetes -- First benchmark walkthrough and Dynamo manifest examples.
- End-to-End Workflow -- Full
init->validate->preflight->profile->resultslifecycle. - Production Deployments -- CI/CD, Kueue, private registries, multi-tenancy, and operations patterns.
- Kubernetes Configuration Reference -- CRD fields, Helm values, and AIPerfJob configuration.
- Monitoring and Troubleshooting -- Watch, debug, logs, and common failure modes.