← PortfolioCompute and infrastructure

Cedana

VMware for GPUs

Team
Neel MasterCEO
Niranjan RavichandraCTO
Founded
2023
Invested
2023
The problem

How do you save a running AI job and pick it up somewhere else?

Click the machine the job is running on to take it away. The job is saved at the driver, moved to another machine and carries on where it stopped.A job runs on one of three machines, and every call it makes to the GPU passes through a layer at the driver on the way down. When its machine is taken back, the GPU memory and the changes to its files are saved as one checkpoint, moved to another machine and put back. The job carries on from where it stopped, while a job with no save has to start again from the beginning.An illustration, not real data.
How saving a program works

means saving a snapshot of a program's state, so it can restart from that point after a failure instead of from scratch. On Linux the standard tool is , short for Checkpoint/Restore in Userspace. It can freeze a running application, write it to disk as a collection of files, and later bring it back from the exact moment it was frozen. Think of a video game's save button, pressed from outside the game.

CRIU can do this because everything it saves, from memory and threads to open files, sockets and pipes, is defined by Linux and looks the same on any hardware. A GPU is another matter. It does things a standard Linux kernel knows nothing about, so CRIU can't save it alone. NVIDIA's cuda-checkpoint utility fills the gap. It locks the driver, lets work already sent to the GPU finish, copies the GPU's memory over to the host and lets go of the GPU. On restore, the memory goes back onto the GPU at its original addresses.

Why bother? Checkpoints give you fault tolerance, and they let a cluster move jobs between machines. They also make cheap capacity usable. Cloud sell spare machines at steep discounts, up to 90% off, on the condition that you hand the machine back when the cloud wants it.

Further reading Application checkpointing (Wikipedia)CRIU (Wikipedia)cuda-checkpoint (NVIDIA on GitHub)Checkpointing CUDA Applications with CRIU (NVIDIA Technical Blog)Spot Instance interruptions (Amazon Web Services)Amazon EC2 Spot Instances (Amazon Web Services)

Why it is hard
  1. i.

    Big jobs break often

    Large training runs spread across hundreds or thousands of GPUs for weeks or months. Meta trained Llama 3 405B on up to 16K H100 GPUs, and in one 54-day stretch of pre-training it logged 466 job interruptions, 419 of them unexpected. Training is synchronous, so a single GPU failure may mean restarting the entire job. Every restart rolls back to the last save.

  2. ii.

    A computer inside the computer

    A GPU is harder to snapshot than a CPU because it is built differently, down to its memory system, its dynamic parallelism and the way its threads synchronise. It is also something of a sealed box. Getting at its state takes low-level access to the driver, and NVIDIA's drivers are proprietary.

  3. iii.

    Every route costs something

    Until recently the only way in was : sit between the program and the GPU, log the calls it makes, and replay them on restore. That adds performance overhead and needs hardware-specific code that is hard to test and maintain. The native route has its own catch. It waits for work already sent to the GPU to finish before a checkpoint completes.

Further reading CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads (arXiv)The Llama 3 Herd of Models (arXiv)Performance of Cedana's GPU Interception (Cedana Docs)Checkpointing CUDA Applications with CRIU (NVIDIA Technical Blog)Cedana vs. CRIUgpu for GPU Checkpoint/Restore (Cedana Docs)

What Cedana is after

Cedana's diagnosis is short: "AI workloads cannot move once running." A job pinned to one machine can't step off a failing node, can't hop onto cheaper capacity and can't be shuffled to fill idle GPUs.

Cedana wants to save, migrate and resume running CPU and GPU workloads across instances and vendors. The goals are to squeeze out idle resources, reach whatever compute is available, and keep work alive through node failures.

Further reading GPU Migration Infrastructure (Cedana)Welcome to Cedana (Cedana Docs)

How they go at it
  1. Step 1: Listen at the driver

    Cedana interposes at the while a process runs. From there it can capture and restore GPU state, even for multi-GPU jobs spread across several machines, and move workloads live. Its optimisations effectively turn it into a virtual machine for NVIDIA GPUs.

  2. Step 2: Earn the overhead back

    Watching every call costs something, so Cedana tries to pay for itself. Seeing every GPU operation lets it merge CUDA graph calls and drop redundant ones, a bit like a just-in-time compiler. And because the CPU often sits idle waiting on the GPU anyway, much of the extra work hides inside that wait. Its own benchmarks show a "minimal overhead of 3%" on raw GPU work.

  3. Step 3: Bring the files along

    A restored process expects its files to be exactly the size it left them. Cedana diffs the container's filesystem and adds that to the checkpoint. For files well over a gigabyte it restores from a Kubernetes volume snapshot instead.

  4. Step 4: Whole clusters, no rewrites

    It works with MPI and , so jobs spread over many GPUs can be snapshotted too, and it plugs into Kubernetes and SLURM without any changes to the code being saved.

Further reading Performance of Cedana's GPU Interception (Cedana Docs)What We Checkpoint (Cedana Docs)GPU Migration Infrastructure (Cedana)

Still open
  • Which layer should do the saving?

    NVIDIA's driver now checkpoints natively, and recent versions add GPU migration. The CRIUgpu researchers point out that this avoids interception's overhead and upkeep. The case for interception is that seeing every call opens up optimisations a native switch can't. Both routes are moving quickly, and it isn't settled which wins where.

  • How do you snapshot a job spread over thousands of GPUs?

    The processes have to checkpoint in a way that is consistent with one another, usually by agreeing on the moment together. If each saves on its own schedule, rollbacks can cascade, in the worst case all the way back to the start, which is known as the . Checkpointing is already one of the major I/O workloads on big systems, so saving more often has a price of its own.

Further reading cuda-checkpoint (NVIDIA on GitHub)CRIUgpu: Transparent Checkpointing of GPU-Accelerated Workloads (arXiv)Cedana vs. CRIUgpu for GPU Checkpoint/Restore (Cedana Docs)Application checkpointing (Wikipedia)

About Cedana

Cedana (YC S23) brings hyperscaler and frontier-lab orchestration capabilities for AI workflows, with GPU live migration as their core capability. Described as "VMware for GPUs," Cedana enables enterprises to orchestrate AI workloads precisely, reliably, and efficiently.

The company builds live migration and checkpoint/restore infrastructure for GPU-heavy workloads, enabling stateful reliability even through catastrophic GPU failures. Their technology interposes at the Driver API to reliably, transparently, and deterministically capture and restore GPU state for multi-GPU workloads.

Performance: Increases cost savings up to 80%, accelerates time to first token 2-10x, and restored cold starts are an order of magnitude faster than native cold starts.

Words used here
Checkpointing
Saving a snapshot of a running program's state so it can later resume from that point.
CRIU
Checkpoint/Restore in Userspace, an open-source Linux tool that freezes a running program to files and restores it.
spot instances
Spare cloud machines sold cheaply on the understanding that the provider can take them back at short notice.
interception
Placing a layer between a program and the GPU so every call to the GPU passes through it and can be recorded.
Driver API
The low-level interface a program uses to talk to NVIDIA's GPU driver.
NCCL
NVIDIA's library for passing data quickly between many GPUs working on one job.
domino effect
A chain of rollbacks in which restoring one process forces others back to ever earlier checkpoints.
Sources