GKE Pod Snapshots: Cold Starts Drop 89%, If Your Nodes Match
Google's benchmarks show a 70B model restoring in 37 seconds instead of minutes. The catch is a hash and a hardware match that silently refuses to restore when either is off.
Key takeaways
- Google's benchmarks show a 70B model restoring in 37 seconds instead of minutes.
- The catch is a hash and a hardware match that silently refuses to restore when either is off.
Google's own benchmarks for GKE Pod snapshots show a 70B parameter model restoring in 37 seconds and an 8B model in 15, against startup latency cuts of up to 89% overall. That fixes the worst part of GPU autoscaling: the minutes a scaled-to-zero inference pod spends reloading weights before it serves a request. What will bite teams is not the mechanism, it is the restore fence: an exact machine series and driver match that silently falls back to a cold boot instead of erroring loudly.
What actually gets checkpointed#
Pod snapshots are checkpoint and restore, not caching. GKE captures process memory, execution threads, CPU registers, and open file descriptors, plus the container's root filesystem, emptyDir volumes, and tmpfs mounts. For GPU workloads, NVIDIA's cuda-checkpoint tool pulls GPU memory into process memory so model weights ride along in the same snapshot. The whole thing is written to Cloud Storage through gVisor, which is why the pod has to be running under GKE Sandbox in the first place. PersistentVolumeClaims, live network connections, and custom routing rules are explicitly left out. On restore, open connections just get re-established the normal way.
That gVisor dependency is the one architectural choice worth internalizing before you plan around this: you are opting into the sandboxed runtime everywhere you want snapshot support, not layering snapshots onto an existing unsandboxed deployment. A snapshot is triggered as its own resource against a running pod, not a flag on the Deployment:
# pod-snapshot-trigger.yaml
apiVersion: podsnapshot.gke.io/v1
kind: PodSnapshotManualTrigger
metadata:
name: inference-pod-snapshot
namespace: model-serving
spec:
targetPod: inference-sandbox-0
$ kubectl apply -f pod-snapshot-trigger.yaml
$ kubectl get podsnapshotmanualtrigger inference-pod-snapshot -n model-serving -w
A new pod requesting the same SandboxTemplate and hashing to the same distilled value gets matched to that snapshot automatically; there is no separate restore command.
The number that matters is 37 seconds, not 89%#
Headline percentages are easy to misread. The 89% figure is a reduction in startup latency across Google's test matrix, but the absolute numbers are what you should plan capacity around: a 70B parameter model comes back in 37 seconds, an 8B model in 15. Compare that against a cold load of a 70B model from object storage into GPU memory, which routinely runs into several minutes. For a scale-to-zero inference service, that gap is the difference between scale to zero being viable and scale to zero meaning a multi-minute queue on every burst.
If your workload never scales past a steady floor of warm replicas, none of this changes your bill. It matters specifically for bursty or idle-heavy GPU inference, the same shape of workload we flagged when HPA and VPA fight over the same signal: the autoscaler that kills a replica during a lull is the one creating the cold start this feature erases.
The restore fence: a hash, a machine series, and two driver versions#
Here is where adoption gets expensive if you skip the fine print. GKE builds what the docs call a distilled Pod spec, a hash derived from the fields that matter for compatibility, and embeds it in the snapshot. A restoring pod has to produce an identical hash. The target node must also run the same machine series and CPU architecture as the node that made the snapshot, N2 to N2 or G2 to G2, never across families, and the gVisor kernel version and GPU driver version baked into the node image have to match what was captured.
Miss any one of those and the pod does not fail loudly. It just starts cold, the normal way, with no error flagging the mismatch as the cause. If your node pools auto-upgrade GPU drivers on a different cadence than you cut new snapshots, you can lose the speedup for weeks without a dashboard pointing at why. The fix is operational: pin the node image and driver version for any pool serving snapshot-restoring workloads, and treat a driver bump as an event that invalidates existing snapshots.
There is a second, looser mode worth knowing: a rootfs-only scope skips the hash check and allows cross-machine-family restores, because it skips process memory, GPU state, and registers too. It is faster to adopt and has none of the hardware fence, but it also gives up the part of the feature that makes GPU cold starts disappear.
Where the GPU cost math actually changes#
The honest case for Pod snapshots is narrow: GPU-backed inference that scales to zero or near-zero between bursts, on a stable node pool with pinned driver and image versions. That combination lets you stop paying for idle accelerator time without eating a multi-minute cold start on every scale-up. It reached general availability on GKE 1.35.3-gke.1234000 and later, so this is not a preview feature worth waiting out.
Outside that lane, the fence cancels the benefit. Mixed-instance node pools that spread G2 and other accelerator families across zones for availability will restore some pods cold without telling you. Clusters on an auto-upgrade channel that bumps GPU drivers on its own schedule need an explicit process to re-snapshot after every driver change, or the feature quietly degrades to doing nothing.
The decision, concretely#
- Do GPU pods scale to zero or near-zero between bursts? Pod snapshots directly address this; the 37-second restore for a 70B model is the number to budget against.
- Is your GPU node pool a single pinned machine series and driver version? If yes, adopt it. If you run mixed G2/A3 pools or auto-upgrading images, fix that first or the hash mismatch will silently erase the win.
- Do you need cross-machine-family flexibility more than GPU-state restore? Use the
rootfs-onlyscope, and budget for a normal cold GPU load on every restore. - Can your platform team own re-snapshotting after driver bumps? If nobody owns that step, the feature degrades on its own within one driver upgrade cycle, with no error to tell you.
The call we'd make#
Adopt GKE Pod snapshots for GPU inference that genuinely scales to zero, on a node pool you keep pinned to one machine series and one driver version on purpose. That is a real fix for the worst part of GPU autoscaling economics, and the 37-second restore number for a 70B model is worth planning around. Do not adopt it on a mixed-instance or auto-upgrading pool and expect the speedup to hold; the distilled Pod spec hash will silently drop you back to a cold start the moment the hardware or driver drifts, and nothing in the restore path will tell you that happened.
Get the DevOps Troubleshooting Cheat Sheet
Subscribe and get our free one-page reference for the errors that eat an afternoon — CrashLoopBackOff, OOMKilled, Terraform state locks, and more — plus new guides as we publish them.
Zed Delta and the Pull Request: Obsolete, or Just Overloaded?
Zed says agents made pull requests obsolete and shipped Delta to replace them. The review unit needs rethinking, but review itself, and the gates around it, stay.
Docker Cloud Sandboxes: Why Agents Need MicroVMs, Not Containers
Docker put AI coding agents in hosted microVMs instead of containers, because a container was never the isolation boundary this job needed.
More from Cloud
Explore more articles in this category
Azure OpenAI's Sweden Central Outage: A Health Check Postmortem
A slow database made a backend service look dead, and the health check that was supposed to protect it killed it faster. The lesson has nothing to do with AI.
AWS Lost a Region for Good: Multi-AZ Is Not Disaster Recovery
AWS says it cannot restore data held only in Bahrain (me-south-1) or in one UAE zone. Multi-AZ gave availability, not recovery, and only cross-region copies survived.
The Cheapest Way to Centralize Logs at Scale
Cutting a log bill is not a procurement exercise. It is four decisions about what you drop at the agent, what you index, how long you keep it, and what you never send at all.
You might have missed
Evergreen posts worth revisiting.