This controller provides Pod Migration—the ability to "move" a running Pod from one node to another while preserving its internal runtime state (including memory, CPU registers, and system sockets). It uses GKE Pod Snapshots to reliably capture the state of the pod and restore it. This feature requires the use of the GKE Sandbox, which is based on the gVisor container runtime.
This controller implements Pod Migration by intercepting eviction requests (e.g. during node drains), triggering manual snapshots with container termination (postCheckpoint: stop), and managing the scheduling of replacement pods. This reduces downtime and prevents progress loss from evictions caused by:
- Node upgrades: When a node is cordoned and drained, the controller captures the state of the pod, preserving in-progress state for Jobs or reducing warm-restore time for stateful database engines.
- VPA evictions: When Vertical Pod Autoscaler evicts pods for resizing, the controller preserves the pod's state prior to eviction so the newly resized pod can resume normal operation faster.
- Other causes: Responds to manual evictions and third-party controllers triggering Pod evictions.
While this controller mitigates the effects of eviction, the migration is not completely transparent and the application will experience:
- A change in Pod identity (different Pod name suffix and IP address).
- Closed network connections that must be re-established by client-side retry logic.
The controller is implemented as a custom Operator consisting of the following core components:
- Eviction Interception Webhook (
/validate-v1-pod-eviction): Handled byEvictionGate. Intercepts Pod eviction API calls. If the Pod is opted-in (pod-migration.gke.io/enabled: "true"), it denies the eviction request with429 (Too Many Requests)to block immediate destruction, and spawns aPodMigrationJob(PMJ) CR to orchestrate the migration flow. - Replacement Webhook (
/mutate-v1-pod-scheduling-gate): Handled byPodGateInjector. Intercepts Pod creation requests. For replacement pods, it injects thegke.io/pod-migration-gatescheduling gate to hold the pod in a pending state until volume detachment is complete, preventing replica collisions. - Status Mutating Webhook (
/mutate-v1-pod-status): Handled byPodStatusMutator. Intercepts Pod status updates. For migrating Job-owned Pods, it overrides the final exit code to137and status toFailed. This allows the Kubernetes Job controller to reschedule the Pod without depleting itsbackoffLimitretry budget (see Job Rescheduling section).
PodMigration(Config) Reconciler: WatchesPodMigrationresources. It automatically provisions the required GKEPodSnapshotStorageConfig(PSSC) and manualPodSnapshotPolicy(PSP) resources withpostCheckpoint: stopin the target namespace.PodMigrationJob(Execution) Reconciler: Coordinates the active migration loop:- Creates a
PodSnapshotManualTrigger(PSMT) targeting the source Pod to initiate snapshotting and container termination. - Monitors the resulting
PodSnapshotuntil it becomesReady(uploaded to GCS). - Transitions the job to
Evictingphase, which signals the Eviction Webhook that it is safe to allow the source Pod deletion. - Waits for the source Pod to be deleted and its volumes to fully detach from the GCE node.
- Transitions the job to
Succeeded, signaling the Pod Gate Reconciler that it is safe to release the replacement Pod.
- Creates a
- Pod Gate Reconciler: Monitors replacement Pods injected with the
gke.io/pod-migration-gatescheduling gate. It holds them in a pending state and releases the gate only after the matchedPodMigrationJobreachesSucceeded(confirming snapshot readiness and volume detachment), then injects the GKE snapshot name annotation to force restore, and removes the scheduling gate to allow startup. It bypasses gate injection for normal scale-up pods.
Before deploying the Pod Migration Controller, ensure your cluster meets the following requirements:
A GKE Standard cluster with the GKE Pod Snapshots addon enabled.
- Minimum GKE Version:
1.36.0-gke.2253000or later (required to support GKE Pod Snapshots with VPA and manual triggers). - Release Channel: Depending on availability in your region, you may need to use the
Rapidrelease channel to obtain a compatible version.
Example cluster creation command (using Rapid channel):
gcloud container clusters create pod-migration-cluster \
--release-channel=rapid \
--workload-pool=<YOUR_PROJECT>.svc.id.goog \
--enable-pod-snapshots \
--zone=<YOUR_ZONE> \
--project=<YOUR_PROJECT>At least one node pool must have gVisor sandboxing enabled (--sandbox=type=gvisor) to run the sandboxed workloads:
gcloud container node-pools create gvisor-pool \
--cluster=pod-migration-cluster \
--sandbox=type=gvisor \
--machine-type=n2-standard-4 \
--num-nodes=2 \
--zone=<YOUR_ZONE> \
--project=<YOUR_PROJECT>cert-manager must be installed to manage TLS certificates for the admission webhooks:
kubectl apply -f https://github.com/cert-manager/cert-manager/releases/download/v1.15.0/cert-manager.yaml
kubectl wait --for=condition=Available --timeout=5m -n cert-manager deployment/cert-manager-webhookFollow these steps to deploy the Pod Migration Controller once the prerequisites are met:
Run the provided setup script to automatically create the GCS bucket, configure IAM permissions for Workload Identity, and create the test Kubernetes Service Account (KSA):
chmod +x patches/setup-storage.sh
./patches/setup-storage.sh [PROJECT_ID] [BUCKET_NAME] [REGION]This script performs the following:
- Creates a GCS bucket
gs://<BUCKET_NAME>with uniform bucket-level access and soft-delete disabled. - Binds IAM roles
roles/storage.objectUserandroles/storage.bucketViewerto all KSAs in thedefaultnamespace and the GKE container engine robot service account. - Creates KSA
pm-test-ksain thedefaultnamespace.
Install the custom resource definitions for the pod migration controller:
kubectl apply -f controller/config/crd/bases/The controller manager orchestrates the migration lifecycle and hosts the admission webhooks.
-
Build and push the controller image:
make -C controller docker-build IMG=<YOUR_REGISTRY>/pod-migration-controller:latest docker push <YOUR_REGISTRY>/pod-migration-controller:latest
-
Configure the GCS Bucket: Edit
controller/podmigration-config.yamlto configure your target GCS bucket for snapshots:spec: storage: location: gs://<YOUR_BUCKET_NAME>/snapshots
-
Deploy the Controller: Apply the manifests by replacing the
<YOUR_CONTROLLER_IMAGE>placeholder:sed 's|<YOUR_CONTROLLER_IMAGE>|<YOUR_REGISTRY>/pod-migration-controller:latest|' controller/deploy.yaml | kubectl apply -f - # Apply the storage config kubectl apply -f controller/podmigration-config.yaml # Verify deployment status kubectl rollout status deployment/pod-migration-controller -n pod-migration-system
Important
Upgrading from Single-Replica to Multi-Replica HA: The controller runs in High Availability (2 replicas) with active/standby Leader Election enabled by default. When upgrading an existing cluster from the legacy un-elected single-replica deployment, scale down and wait for the old pod to fully terminate before applying the HA manifests:
kubectl scale deployment/pod-migration-controller -n pod-migration-system --replicas=0
kubectl wait --for=delete pod -l app=pod-migration-controller -n pod-migration-system --timeout=60s
sed 's|<YOUR_CONTROLLER_IMAGE>|<YOUR_REGISTRY>/pod-migration-controller:latest|' controller/deploy.yaml | kubectl apply -f -To reject incompatible workloads (e.g. BEAM/fsnotify) at admission time:
kubectl apply -f patches/gke-pod-snapshot-admission-webhook.yamlEvery workload in this matrix has been verified through real E2E eviction migrations on a GKE Standard node pool.
Note
- Job Rescheduling: Workloads of type
Job(e.g.,go,node) are supported for resilient rescheduling using the Status Mutating Webhook. - NATS Caveat:
natsis marked as flaky due to a transient runtime-level deadlock during checkpointing; see troubleshooting notes below.
| Application | Category | Verdict | Required Workarounds & Bypass Configurations |
|---|---|---|---|
| node | app / Job | ✅ SURVIVED | Verified both as StatefulSet and Job (rescheduled via webhook). Out-of-the-box support. |
| go | app / Job | ✅ SURVIVED | Verified both as StatefulSet and Job (rescheduled via webhook). Out-of-the-box support. |
| redis | datastore | ✅ SURVIVED | Pure in-memory key-value store. Very fast migration. |
| valkey | datastore | ✅ SURVIVED | Pure in-memory key-value store. Very fast migration. |
| mysql | datastore | ✅ SURVIVED | InnoDB AIO Bypass: Disable native async I/O (--innodb_use_native_aio=OFF) to avoid seccomp blocks on host io_uring. |
| mariadb | datastore | ✅ SURVIVED | InnoDB AIO Bypass: Disable native async I/O (--innodb_use_native_aio=OFF) to avoid seccomp blocks on host io_uring. |
| memcached | datastore | ✅ SURVIVED | Pure in-memory cache blocks restored. |
| dragonfly | datastore | ✅ SURVIVED | epoll Bypass: Requires epoll forcing flag (--force_epoll). Verified with redis-client. |
| vault | secrets | ✅ SURVIVED | Dev-mode secrets state restored. |
| consul | coordination | ✅ SURVIVED | Dev-mode memory key-value databases survive. |
| etcd | coordination | ✅ SURVIVED | BoltDB storage writes successfully restored. |
| nats | streaming | Deadlock Flake / Retry Caveat: Memory jetstream offsets restored. Encountered runtime-level deadlock on 1st run. Succeeded on retry. | |
| zookeeper | coordination | ✅ SURVIVED | emptyDir Path Redirect: Redirect ZOO_DATA_DIR away from Kubelet emptyDir mounts (e.g. to /tmp/zookeeper) to prevent walk errors. |
| kafka (KRaft) | streaming | ✅ SURVIVED | JVM Metrics Bypass: Inject environment variable KAFKA_OPTS="-XX:-UseContainerSupport" to avoid cgroups mismatch crashes on target nodes. |
| postgres | datastore | ✅ SURVIVED | Works out-of-the-box (uses guest POSIX shared memory). Requires setting PGDATA to container local directories. |
| minio | datastore | ✅ SURVIVED | Redirect storage paths to container writable layers to avoid emptyDir mount walk failures. |
| nginx | proxy | ✅ SERVED | Stateless proxies restore and handle reconnected traffic. |
| haproxy | proxy | ✅ SERVED | Stateless proxies restore and handle reconnected traffic. |
| traefik | proxy | ✅ SERVED | Stateless routers restore and handle reconnected traffic. |
| caddy | proxy | ✅ SERVED | Stateless routers restore and handle reconnected traffic. |
| python (HTTP) | app | ✅ SERVED | Stateless python workers survive. |
| mongodb | datastore | ❌ FAILED | WiredTiger storage engine locks and blocks on sandboxed io_uring seccomp. |
| cassandra | datastore | ❌ FAILED | Large JVM heaps mismatch host cgroups descriptors post-restore. |
| cockroachdb | datastore | ❌ FAILED | Raft synchronization timeouts and socket reset crashes post-restore. |
| clickhouse | datastore | ❌ FAILED | Columnar block datastore sync locks and file descriptor leaks. |
| rabbitmq / couchdb | streaming | ❌ FAILED | Erlang BEAM runtime epoll and green thread scheduler structures cannot be serialized. |
| prometheus | monitoring | ❌ REFUSED | Active WAL memory mappings exceed serialization limits under both runtimes. |
| elasticsearch | search | ❌ REFUSED | Heavy fsnotify directory watches cannot be serialized. |
To support resilient rescheduling of Jobs during migration, the system uses a Status Mutating Webhook combined with Kubernetes Job Pod Failure Policy.
A migrating Job Pod can terminate for two distinct reasons, and each needs its own podFailurePolicy match. Cover both, or the Job will burn backoffLimit retries on events that are not application failures.
- When a migrating Job Pod is terminated, the
pod-migration-controller's status mutating webhook (/mutate-v1-pod-status) intercepts the status update. - It mutates the Pod phase to
Failedand sets the container exit code to137. - The Job controller interprets this exit code according to the Job's
podFailurePolicy. - If configured correctly, the Job controller ignores this failure and recreates the Pod (which then restores from the snapshot) without counting it against the Job's
backoffLimitretry budget.
When a replacement Pod fails to restore from its snapshot, the controller fails the migration and deletes the Pod so its Job recreates it cold. This is a recovery action, not an application failure, but it is not covered by the rule above:
- The status webhook only rewrites exit codes on Pods transitioning to
Succeededduring an activeSnapshotting/Evictingmigration. The crashed replacement Pod is in neither state, so it is never touched and never receives the synthetic137. - A container the runtime failed to start terminates with
reason: StartErrorand exit code128.
Without a match for 128, every cold-start fallback silently consumes one of the Job's backoffLimit retries.
Important
Precondition: onExitCodes matches the exit code alone — it cannot additionally require reason: StartError. An Ignore match on 128 therefore assumes your application never exits 128 itself. If it can, leave 128 out and treat these as application failures; you will spend one retry per fallback instead.
Why not a DisruptionTarget condition rule?
Kubernetes only adds the DisruptionTarget Pod condition for evictions via the Eviction API, kubelet node-pressure and graceful shutdown, the taint manager, PodGC, and scheduler preemption. The fallback issues a direct DELETE, so the condition is never set and such a rule would never fire.
Routing the fallback through the Eviction API to obtain the condition would make it PDB-aware — and the crashed Pod is not Ready, so under the default unhealthyPodEvictionPolicy: IfHealthyBudget the eviction can be refused, wedging the exact Pod the fallback exists to remove. Separately, onPodConditions matches on type and status but not reason, so the rule would also swallow preemptions and node-pressure evictions. The exit code is both narrower and more reliable here.
To enable this, your Job manifest must:
- Enable migration via labels.
- Configure
podFailurePolicytoIgnoreexit codes137(migration eviction) and128(restore-crash fallback).
Here is an example snippet (from verification-suite/manifests/pm-go-job.yaml):
apiVersion: batch/v1
kind: Job
metadata:
name: pm-go-job
spec:
backoffLimit: 4
podFailurePolicy:
rules:
- action: Ignore
onExitCodes:
operator: In
# 137: pod evicted by migration (synthesized by the status webhook)
# 128: container start failure -> restore-crash cold-start fallback
values: [137, 128]
template:
metadata:
labels:
pod-migration.gke.io/enabled: "true"
spec:
runtimeClassName: gvisor
restartPolicy: Never
# ... rest of the specTo migrate your workload using this controller:
Add the label pod-migration.gke.io/enabled: "true" to your workload Pod template and ensure the Pod uses the gvisor runtime.
The eviction webhook only migrates Pods whose runtime class is listed in the controller flag --migratable-runtime-classes (default gvisor; the token @default matches Pods that set no runtimeClassName). Only list a class that your installed snapshot engine supports: GKE Pod Snapshots supports gvisor.
Example (Deployment):
apiVersion: apps/v1
kind: Deployment
metadata:
name: my-app
spec:
replicas: 1
selector:
matchLabels:
app: my-app
template:
metadata:
labels:
app: my-app
pod-migration.gke.io/enabled: "true" # Enables migration orchestration
spec:
runtimeClassName: gvisor # Required for GKE Pod Snapshots
containers:
- name: app
image: my-app-image:latestThe controller automatically intercepts standard Kubernetes evictions and orchestrates the stateful migration. You can trigger this manually (e.g. for testing node upgrades) by draining the node the Pod is running on:
- Find the node the Pod is running on:
kubectl get pod -l app=my-app -o wide
- Drain the node to trigger eviction:
kubectl drain <node-name> --ignore-daemonsets --delete-emptydir-data --force --grace-period=30
- Uncordon the node once the migration starts to make it available again for rescheduling:
kubectl uncordon <node-name>
This repository contains production-ready YAML templates for trying out pod migration on your workloads under the verification-suite/manifests/ directory.
You can run automated E2E pod migration verification for any of the 21 pre-configured applications using the driver script verification-suite/run_app_validation.sh:
# Run validation on Valkey
./verification-suite/run_app_validation.sh valkey
# Run validation on MySQL
./verification-suite/run_app_validation.sh mysqlWhen running ./verification-suite/run_app_validation.sh valkey, you should expect logs similar to:
[*] Deploying Valkey StatefulSet...
statefulset.apps/pm-valkey created
service/pm-valkey-service created
[*] Waiting for Valkey pod to be Ready...
pod/pm-valkey-0 condition met
[*] Seeding state in Valkey: migkey -> valkey-nonce-1718900000
OK
[*] Pod is running on node: gke-pod-migration-cluster-gvisor-pool-abcdef-1234
[*] Draining node gke-pod-migration-cluster-gvisor-pool-abcdef-1234...
node/gke-pod-migration-cluster-gvisor-pool-abcdef-1234 cordoned
evicting pod pm-system/pm-controller-manager-...
evicting pod default/pm-valkey-0
node/gke-pod-migration-cluster-gvisor-pool-abcdef-1234 drained
[*] Restoring node gke-pod-migration-cluster-gvisor-pool-abcdef-1234 (uncordon)...
node/gke-pod-migration-cluster-gvisor-pool-abcdef-1234 uncordoned
[*] Waiting for restored Valkey pod to be Ready...
pod/pm-valkey-0 condition met
[*] Verifying state...
[+] Retrieved value: valkey-nonce-1718900000
[SUCCESS] Valkey E2E Live Migration Succeeded. State survived!
PodMigration resources use the podmigration.gke.io/storage-cleanup finalizer to ensure that cluster-scoped PodSnapshotStorageConfig (PSSC) and namespaced PodSnapshotPolicy (PSP) resources are cleanly deleted when a PodMigration CR is deleted. The finalizer also verifies whether any active migrations (PodMigrationJob resources in non-terminal phases) are running in the namespace, postponing deletion until in-flight jobs reach a terminal state (Succeeded, SucceededWithoutRestore, or Failed).
When undeploying the controller, PodMigration custom resources must be deleted while the controller is still running so it can process the finalizer:
# make undeploy handles this automatically by deleting PodMigration CRs first:
make -C controller undeployIf the controller is deleted before the custom resources, PodMigration objects will wedge in Terminating and block CRD deletion. Similarly, if the controller is downgraded or rolled back to a pre-#35 image that does not support the finalizer, any PodMigration deleted during that time will wedge in Terminating. To manually unwedge:
kubectl patch podmigration <name> -n <namespace> --type=json -p='[{"op": "remove", "path": "/metadata/finalizers"}]'GKE enforces a ValidatingAdmissionPolicy (gke-pod-snapshot-validating-admission-policy) that prevents manual edits to podsnapshots by anyone other than the GKE snapshot controller and agent. This blocks users from manually removing finalizers from stuck podsnapshots (e.g. when GCS upload fails or during test resets).
To override this check and perform cleanup:
- Disable the validation actions temporarily by patching the binding to "Audit" instead of "Deny":
kubectl patch validatingadmissionpolicybinding gke-pod-snapshot-vap-binding \ --type=json -p='[{"op": "replace", "path": "/spec/validationActions", "value": ["Audit"]}]' - Remove finalizers and delete the stuck snapshots:
kubectl get podsnapshots -o json | jq -r '.items[].metadata.name' | xargs -I {} kubectl patch podsnapshot {} --type=json -p='[{"op": "remove", "path": "/metadata/finalizers"}]' || true kubectl delete podsnapshots --all --timeout=15s
- Restore the validation actions back to "Deny, Audit":
kubectl patch validatingadmissionpolicybinding gke-pod-snapshot-vap-binding \ --type=json -p='[{"op": "replace", "path": "/spec/validationActions", "value": ["Deny", "Audit"]}]'
When migrating single-replica workloads (such as a StatefulSet with replicas: 1 or a bare Pod) configured with a strict PodDisruptionBudget (minAvailable: 1 or maxUnavailable: 0), Kubernetes PDB admission will prevent eviction of the origin pod because disruptions are not allowed while the single replica is running:
- PDB-Safe Eviction Fallback: The controller respects Kubernetes PDB constraints by issuing eviction requests through the official Kubernetes eviction subresource (
policy/v1 Eviction) and retrying on429 Too Many Requests/409 Conflict. - Timeout & Churn Protection: If the PDB budget does not recover within the 10-minute active migration timeout, the PMJ concludes with status
SucceededWithoutRestore(Reason:PDBEvictionTimeout), and the controller marks the pod withpod-migration.gke.io/pdb-eviction-timeout: "true"to prevent repeated snapshot loops during ongoing drain retries. - Operator Intervention & Re-arming Migration: Draining nodes hosting
minAvailable=1single-replica workloads requires operator intervention (e.g. temporarily updating or relaxing the PDB budget, scaling up the workload, or deleting the pod). Once the PDB is relaxed, to re-arm live migration on the existing pod instance without recreating it, remove the timeout annotation:kubectl annotate pod <pod-name> pod-migration.gke.io/pdb-eviction-timeout-
The controller registers Prometheus metrics on the standard controller-runtime metrics endpoint (:8080/metrics by default):
| Metric | Type | Labels | Description |
|---|---|---|---|
pod_migration_active |
Gauge | None | Number of currently active pod migrations in flight. |
pod_migration_outcomes_total |
Counter | outcome (succeeded, failed, timeout, fallback, succeeded_without_restore) |
Total completed migrations by outcome. |
pod_migration_phase_duration_seconds |
Histogram | phase (pending, snapshotting, evicting, restoring) |
Latency distribution of individual migration phases. |
pod_migration_restore_crash_fallback_total |
Counter | None | Total cold-start fallbacks triggered by matched gVisor/OCI restore crashes. |
pod_migration_restore_crash_unmatched_total |
Counter | None | Total unrecognised StartError restore-pod crashes with no known signature. |
When running in multi-replica HA mode with leader election enabled:
- The
/metricsendpoint is served by all controller replicas. - Only the leader replica reconciles migrations and updates
pod_migration_activein memory; standby replicas servepod_migration_active == 0. - Prometheus scrape jobs should scrape per-pod endpoints and use
max(pod_migration_active)when querying or alerting on active in-flight migrations across the cluster. - For counter and histogram metrics, use
sum(rate(pod_migration_outcomes_total[5m])) by (outcome)andsum(rate(pod_migration_phase_duration_seconds_sum[5m])) by (phase) / sum(rate(pod_migration_phase_duration_seconds_count[5m])) by (phase).