Cedana
VMware for GPUs
How do you save a running AI job and pick it up somewhere else?
Saving a snapshot of a running program's state so it can later resume from that point. means saving a snapshot of a program's state, so it can restart from that point after a failure instead of from scratch. On Linux the standard tool is Checkpoint/Restore in Userspace, an open-source Linux tool that freezes a running program to files and restores it., short for Checkpoint/Restore in Userspace. It can freeze a running application, write it to disk as a collection of files, and later bring it back from the exact moment it was frozen. Think of a video game's save button, pressed from outside the game.
CRIU can do this because everything it saves, from memory and threads to open files, sockets and pipes, is defined by Linux and looks the same on any hardware. A GPU is another matter. It does things a standard Linux kernel knows nothing about, so CRIU can't save it alone. NVIDIA's cuda-checkpoint utility fills the gap. It locks the driver, lets work already sent to the GPU finish, copies the GPU's memory over to the host and lets go of the GPU. On restore, the memory goes back onto the GPU at its original addresses.
Why bother? Checkpoints give you fault tolerance, and they let a cluster move jobs between machines. They also make cheap capacity usable. Cloud Spare cloud machines sold cheaply on the understanding that the provider can take them back at short notice. sell spare machines at steep discounts, up to 90% off, on the condition that you hand the machine back when the cloud wants it.
Further reading Application checkpointing (Wikipedia)CRIU (Wikipedia)cuda-checkpoint (NVIDIA on GitHub)Checkpointing CUDA Applications with CRIU (NVIDIA Technical Blog)Spot Instance interruptions (Amazon Web Services)Amazon EC2 Spot Instances (Amazon Web Services)
- i.
Big jobs break often
Large training runs spread across hundreds or thousands of GPUs for weeks or months. Meta trained Llama 3 405B on up to 16K H100 GPUs, and in one 54-day stretch of pre-training it logged 466 job interruptions, 419 of them unexpected. Training is synchronous, so a single GPU failure may mean restarting the entire job. Every restart rolls back to the last save.
- ii.
A computer inside the computer
A GPU is harder to snapshot than a CPU because it is built differently, down to its memory system, its dynamic parallelism and the way its threads synchronise. It is also something of a sealed box. Getting at its state takes low-level access to the driver, and NVIDIA's drivers are proprietary.
- iii.
Every route costs something
Until recently the only way in was Placing a layer between a program and the GPU so every call to the GPU passes through it and can be recorded.: sit between the program and the GPU, log the calls it makes, and replay them on restore. That adds performance overhead and needs hardware-specific code that is hard to test and maintain. The native route has its own catch. It waits for work already sent to the GPU to finish before a checkpoint completes.
Further reading CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads (arXiv)The Llama 3 Herd of Models (arXiv)Performance of Cedana's GPU Interception (Cedana Docs)Checkpointing CUDA Applications with CRIU (NVIDIA Technical Blog)Cedana vs. CRIUgpu for GPU Checkpoint/Restore (Cedana Docs)
Cedana's diagnosis is short: "AI workloads cannot move once running." A job pinned to one machine can't step off a failing node, can't hop onto cheaper capacity and can't be shuffled to fill idle GPUs.
Cedana wants to save, migrate and resume running CPU and GPU workloads across instances and vendors. The goals are to squeeze out idle resources, reach whatever compute is available, and keep work alive through node failures.
Further reading GPU Migration Infrastructure (Cedana)Welcome to Cedana (Cedana Docs)
- Step 1: Listen at the driver
Cedana interposes at the The low-level interface a program uses to talk to NVIDIA's GPU driver. while a process runs. From there it can capture and restore GPU state, even for multi-GPU jobs spread across several machines, and move workloads live. Its optimisations effectively turn it into a virtual machine for NVIDIA GPUs.
- Step 2: Earn the overhead back
Watching every call costs something, so Cedana tries to pay for itself. Seeing every GPU operation lets it merge CUDA graph calls and drop redundant ones, a bit like a just-in-time compiler. And because the CPU often sits idle waiting on the GPU anyway, much of the extra work hides inside that wait. Its own benchmarks show a "minimal overhead of 3%" on raw GPU work.
- Step 3: Bring the files along
A restored process expects its files to be exactly the size it left them. Cedana diffs the container's filesystem and adds that to the checkpoint. For files well over a gigabyte it restores from a Kubernetes volume snapshot instead.
- Step 4: Whole clusters, no rewrites
It works with MPI and NVIDIA's library for passing data quickly between many GPUs working on one job., so jobs spread over many GPUs can be snapshotted too, and it plugs into Kubernetes and SLURM without any changes to the code being saved.
Further reading Performance of Cedana's GPU Interception (Cedana Docs)What We Checkpoint (Cedana Docs)GPU Migration Infrastructure (Cedana)
Which layer should do the saving?
NVIDIA's driver now checkpoints natively, and recent versions add GPU migration. The CRIUgpu researchers point out that this avoids interception's overhead and upkeep. The case for interception is that seeing every call opens up optimisations a native switch can't. Both routes are moving quickly, and it isn't settled which wins where.
How do you snapshot a job spread over thousands of GPUs?
The processes have to checkpoint in a way that is consistent with one another, usually by agreeing on the moment together. If each saves on its own schedule, rollbacks can cascade, in the worst case all the way back to the start, which is known as the A chain of rollbacks in which restoring one process forces others back to ever earlier checkpoints.. Checkpointing is already one of the major I/O workloads on big systems, so saving more often has a price of its own.
Further reading cuda-checkpoint (NVIDIA on GitHub)CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads (arXiv)Cedana vs. CRIUgpu for GPU Checkpoint/Restore (Cedana Docs)Application checkpointing (Wikipedia)
Cedana (YC S23) brings hyperscaler and frontier-lab orchestration capabilities for AI workflows, with GPU live migration as their core capability. Described as "VMware for GPUs," Cedana enables enterprises to orchestrate AI workloads precisely, reliably, and efficiently.
The company builds live migration and checkpoint/restore infrastructure for GPU-heavy workloads, enabling stateful reliability even through catastrophic GPU failures. Their technology interposes at the Driver API to reliably, transparently, and deterministically capture and restore GPU state for multi-GPU workloads.
Performance: Increases cost savings up to 80%, accelerates time to first token 2-10x, and restored cold starts are an order of magnitude faster than native cold starts.
- Checkpointing
- Saving a snapshot of a running program's state so it can later resume from that point.
- CRIU
- Checkpoint/Restore in Userspace, an open-source Linux tool that freezes a running program to files and restores it.
- spot instances
- Spare cloud machines sold cheaply on the understanding that the provider can take them back at short notice.
- interception
- Placing a layer between a program and the GPU so every call to the GPU passes through it and can be recorded.
- Driver API
- The low-level interface a program uses to talk to NVIDIA's GPU driver.
- NCCL
- NVIDIA's library for passing data quickly between many GPUs working on one job.
- domino effect
- A chain of rollbacks in which restoring one process forces others back to ever earlier checkpoints.
- 1CRIU · Wikipedia
- 2Application checkpointing · Wikipedia
- 3Checkpointing CUDA Applications with CRIU · NVIDIA Technical Blog
- 4cuda-checkpoint · NVIDIA on GitHub
- 5CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads · arXiv
- 6The Llama 3 Herd of Models · arXiv
- 7Spot Instance interruptions · Amazon Web Services
- 8Amazon EC2 Spot Instances · Amazon Web Services
- 9GPU Migration Infrastructure · Cedana
- 10Welcome to Cedana · Cedana Docs
- 11Performance of Cedana's GPU Interception · Cedana Docs
- 12What We Checkpoint · Cedana Docs
- 13Cedana vs. CRIUgpu for GPU Checkpoint/Restore · Cedana Docs