GPU TRAINING CONTROL PLANE

The control plane
for GPU training.

Find available GPUs and keep your PyTorch, JAX, or Hugging Face training on them — VaultLayer checkpoints continuously and auto-resumes on fresh hardware through every interruption, so a reclaimed GPU sets you back minutes, not the whole run. Affordable capacity, or your own cloud.

Bring your own cloud — AWS, Azure, GCP, or any provider.

rahuljain@vaultlayer.cloud

$ pip install vaultlayer && vaultlayer init
$ vl run train.py
✓ checkpoint synced
⚠ hardware interrupted at step 2,140
→ resumed from last checkpoint on fresh hardware
✓ run complete — adapter saved
3,244
training jobs run
717
GPU-hours run
HOW IT WORKS

Run your job. We handle the rest.

your training job
python train.py
VaultLayer control plane
checkpoint · monitor · recover
job resumes
after failures
01

Sign up and run vaultlayer init.

02

vl run train.py — no SDK, no decorators, no YAML.

03

On preemption or failure, resumes from the last checkpoint automatically.

WHY

Find a GPU. Finish the run.

Available capacity, found for you.

VaultLayer finds available GPU capacity across the market, provisions it, and runs your job — no hunting for a free region, no managing instances. On affordable GPUs.

Interruptions don't lose your work.

VaultLayer checkpoints your training continuously. When a GPU is preempted or crashes, it brings up fresh hardware and resumes from the last checkpoint — not from the start.

Or bring your own cloud.

Run on your own account — AWS, Azure, GCP, or any provider — with no VaultLayer per-run charge. Your credits and contracts apply.

PROOF

One real run, from our production database.

28.37
GPU-hours of work
8.0h
longest any single machine carried it
10
resumes from checkpoint
$111.56
all-in · completed

The run needed more hours than any machine under it lived. It finished anyway — because the run doesn't depend on any one machine surviving.

one bar = the whole run · each break = a machine turning over · last segment = completed

72B QLoRA fine-tune, one H200-class GPU at a time. The machines under it turned over 10 times — instance recycles, a full disk, a stall — and the job resumed from its last checkpoint every time. Measured, not modeled.

VaultLayer managed — what it actually charged$111.56
Same hours at on-demand list price, same GPU class$217.26

The cheap raw capacity underneath isn't an alternative for a 28-hour job — this run's machines turned over 10 times along the way. The system resumed from checkpoint every time. That's what you're paying for.

"The babysitting tax. … I'd lose 20–30 minutes per session to wrapper work around the actual training … The fine-tune finished cleanly — adapter saved, no leftover pod to clean up, no manual sync, no babysitting."
Amit Pal · CTO, Vettable — design partner

Fails closed: jobs error hard if credentials are missing — no silent fallback

Checkpoint-resume validated end-to-end across our routing pool

Managed checkpoint credentials are job-scoped and short-lived

THE MATH

Why finishing beats restarting.

In our cost model, restarting from scratch compounds with run length until raw spot costs more than on-demand around day three — while checkpoint-and-resume grows roughly linearly, its effective hourly rate staying stable. Assumptions are labeled on the chart; the full model lives in our guides.

FAQ
How does pricing work?

Managed runs are billed at the provider's rate plus a flat percentage platform fee shown before the run starts. BYOC is a monthly platform contract — no VaultLayer per-run charges; your cloud bills you directly for compute.

What can VaultLayer access?

Only what's needed to run your job: the credentials you provide, used to provision and tear down instances. It fails closed if none exist — no silent fallback.

How does checkpoint & resume work with my trainer?

Checkpoints save on your cadence and sync continuously. On interruption, the job restarts from the last checkpoint on fresh hardware. Integration is usually auto-inserted on the first run.

What workloads are supported?

Training and fine-tuning — PyTorch, JAX, and Hugging Face. If it runs with python train.py, it works. Not real-time inference serving.

How do I get access?

Sign up, install the CLI with pip install vaultlayer, and submit your first job with vl run train.py. Or email the founder directly — rahuljain@vaultlayer.cloud.

Leave the babysitting to us.

Bring your training job. It finishes even when the hardware doesn't.

Sign up

Or email the founder directly — rahuljain@vaultlayer.cloud