SWE-perf provides a suite of standalone Docker containers that emulate the behavior of autonomous coding agents (like SWE-agent or OpenHands). This suite allows teams to evaluate infrastructure performance under the heavy, bursty load of AI agents, without actually needing to run an expensive LLM in the loop.
- The Image Suite
- Getting the Images
- Image Properties & Usage
- Building / Regenerating Images (Advanced)
- Cluster Infrastructure & Benchmarking
- Testing
The core value of SWE-perf is the container suite itself. Instead of hitting an OpenAI or Gemini API, the containers use an embedded Log-Normal Replay Engine. This engine:
- Takes a pre-recorded agent trajectory (commands run during a SWE-bench task).
- Simulates character-by-character shell typing using
pexpectagainst a bash session. - Synthesizes realistic LLM latency (think time) between commands using a log-normal distribution derived from real-world agent trajectories (averaging ~15.6s).
Two Operation Modes:
- On-Demand Runner (Singleton Image): A single, unified Docker image (
sweperf:universal). When it starts, it dynamically uses pre-extracted setup scripts to fetch dependencies and builds the SWE-bench environment for the specific task at runtime. This simulates a normal agentic workload where the agent operates on arbitrary repositories on the fly. - Pre-baked Images: A suite of 500+ separate images (one for each SWE-bench task). This simulates a Reinforcement Learning (RL) workload where the environment is thoroughly known, cached, and pre-baked (because an RL agent repeatedly interacts with the same known repository during training).
These standalone containers can be deployed in any environment to simulate realistic AI agent workloads natively.
If you want to use the pre-built, production-ready SWE-perf image suite within your own Google Cloud projects, you do not need to generate them yourself. Instead, synchronize a copy of the official repository directly into your own Artifact Registry. This syncs over both the singleton On-Demand image and the Pre-baked suite.
The most efficient way to achieve this is via a server-to-server copy using gcrane (a Google-maintained CLI for container registries). This bypasses downloading hundreds of gigabytes locally and takes only seconds to copy all images natively.
# 1. Authenticate to Google Cloud
gcloud auth login
gcloud auth configure-docker us-central1-docker.pkg.dev
# 2. Trigger the server-to-server copy into your destination Artifact Registry
./sweperf copy-images \
us-central1-docker.pkg.dev/<YOUR_PROJECT>/<YOUR_REPO_NAME>
# 3. Generate your local image index to point to your freshly-copied repo
./sweperf generate-image-list \
us-central1-docker.pkg.dev/<YOUR_PROJECT>/<YOUR_REPO_NAME>By explicitly copying the images into your own project and generating the local manifests, your GKE clusters and VMs gain native, frictionless access without having to navigate cross-project IAM restrictions or service account key sharing.
Here is an example Pod spec using the singleton On-Demand image:
apiVersion: v1
kind: Pod
metadata:
name: sweperf-agent-test
spec:
containers:
- name: agent
# Replace with your project and repo
image: us-central1-docker.pkg.dev/<YOUR_PROJECT>/<YOUR_REPO_NAME>/sweperf:universal
env:
- name: USE_GCP_CACHE
value: "1"
- name: INSTANCE_ID
value: "astropy__astropy-14365"
- name: WAIT_FOR_CLAIM_FILE
value: "/etc/podinfo/labels"
volumeMounts:
- name: podinfo
mountPath: /etc/podinfo
volumes:
- name: podinfo
downwardAPI:
items:
- path: "labels"
fieldRef:
fieldPath: metadata.labels
restartPolicy: NeverOnce you have access to the images, you can deploy them directly. There are several useful dials baked into the container, controlled via environment variables.
Agent Latency If you want to evaluate maximum cluster churn without being bottlenecked by simulated LLM think time, you can override the distribution parameters.
LLM_LATENCY_MU: (Default:2.0414). Set to a large negative number like"-5"to floor out the latency for extreme churn testing (e.g. constant 0.5s delay).LLM_LATENCY_SIGMA: (Default:0.8674)LLM_LATENCY_MIN: (Default:0.5)
Network & Sandbox Blocking
WAIT_FOR_CLAIM_FILE: To use the generated images seamlessly with Agent Sandbox's warm pools, set this to a file path containing downward API labels (e.g./etc/podinfo/labels). The container will pause all execution (including repository setup) until it finds theagents.x-k8s.io/sandbox-idlabel inside that file.WAIT_FOR_START_PORT: If you are not using Agent Sandbox but still want to deploy pods in a "warm" paused state, set this to a port number (e.g.,8080). The container will stand up a simple TCP server and pause execution until it receives an HTTP GET/startrequest on that port.
Caching & On-Demand Options (Universal Image Only) When running the On-Demand Universal Image (e.g., downloading repositories and packages dynamically), these variables control how the system clones repositories and downloads pip dependencies.
INSTANCE_ID: (Required for On-Demand) The specific SWE-bench task name the container should evaluate (e.g.,astropy__astropy-14365). This dictates the test files, the package commit, and the exact setup script run.DO_GIT_PULL: (Default:false) Selects whethergit pullupdates the repository from upstream, depending on your container repository caching strategy.USE_GCP_CACHE: (Default:"0") If set to"1", the pod entrypoint will securely ping the internal GCP metadata server to fetch an ephemeral OAuth token for Google Cloud Artifact Registry (used to authenticate a PyPI pull-through proxy).PIP_INDEX_URL_TEMPLATE: A PyPI proxy URL string containing a{TOKEN}placeholder template (e.g.,https://oauth2accesstoken:{TOKEN}@us-central1-python.pkg.dev/my-project/my-repo/simple/). This relies onUSE_GCP_CACHEdiscovering a token and substituting it securely into the URL.PIP_TRUSTED_HOST: Explicitly specifies the trusted private host, circumventing SSL certificate resolution if a private proxy isn't on a recognized root domain.
Note: For out-of-the-box cluster orchestration, we recommend using the ./sweperf create-caches command to spin up Python pull-through proxies automatically, and passing its resulting URL linearly to ./sweperf run-universal 2 300 <URL>. This instructs the submitter bash loop to automatically plumb all of the above options into the test pod environment variables.
If you have made edits to the replay engine or script injector, you will need to re-generate the image suite from scratch.
First, build the generic base environment (sweb.base.x86_64:latest):
./sweperf build-sweb-baseThen extract setup bash scripts, and build the single On-Demand Docker image (sweperf:universal):
./sweperf generate-universal [--image-prefix <PREFIX>]To generate the legacy 500+ Docker images, push them to a registry, and optionally pre-pull them onto nodes:
./sweperf generate-images --project <PROJECT_ID> --region <REGION> --repo <REPO_NAME> [--limit <LIMIT>] [--run <RUN_NAME>]While the generated images can be used anywhere, SWE-perf also includes scripts to provision a high-density GKE cluster, set up caches, and run structured load tests for both On-Demand and Pre-baked workloads.
Create a GKE cluster with a high-density footprint (--default-max-pods-per-node=256). This command will also automatically provision an Artifact Registry repository and synchronize the SWE-perf image suite into it.
./sweperf create-cluster <PROJECT_ID> <REGION> <CLUSTER_NAME> <REPO_NAME>Since the On-Demand Runner dynamically builds the environment for each task at runtime, it relies heavily on fetching dependencies (like pip install). To avoid getting rate-limited or bogged down by network latency during high-density tests, deploy a PyPI proxy cache:
./sweperf create-caches <PROJECT_ID> <REGION> <PYPI_REPO>To avoid network throttling and excessive disk I/O when spinning up hundreds of distinct pods simultaneously, cache the pre-baked test images onto your nodes beforehand:
./sweperf prepull-imagesThe submitter commands spin up a stateless submitter alongside a metrics collector. It queries the Kubernetes API and maintains a strict concurrent pod limit over a user-defined time window, automatically backfilling pods as they complete.
On-Demand Benchmark
./sweperf run-universal <DURATION_SECONDS> <CONCURRENCY> [PYPI_CACHE_URL]
# Example (20 minutes with 512 active pods, utilizing the PyPI cache):
# ./sweperf run-universal 1200 512 us-central1-python.pkg.dev/my-project/my-pypi-cachePre-baked Benchmark
./sweperf run-benchmark <DURATION_SECONDS> <CONCURRENCY>To run the local unit tests that verify the replay.py execution within a dummy Docker container using pexpect:
python3 -m unittest discover -s testsThis project is licensed under the Apache 2.0 License.
We welcome contributions! Please see docs/contributing.md for more information.
We follow Google's Open Source Community Guidelines.
This is not an officially supported Google product.
This project is not eligible for the Google Open Source Software Vulnerability Rewards Program.