Find available GPUs and keep your PyTorch, JAX, or Hugging Face training on them — VaultLayer checkpoints continuously and auto-resumes on fresh hardware through every interruption, so a reclaimed GPU sets you back minutes, not the whole run. Affordable capacity, or your own cloud.
Bring your own cloud — AWS, Azure, GCP, or any provider.
rahuljain@vaultlayer.cloud
Sign up and run vaultlayer init.
vl run train.py — no SDK, no decorators, no YAML.
On preemption or failure, resumes from the last checkpoint automatically.
VaultLayer finds available GPU capacity across the market, provisions it, and runs your job — no hunting for a free region, no managing instances. On affordable GPUs.
VaultLayer checkpoints your training continuously. When a GPU is preempted or crashes, it brings up fresh hardware and resumes from the last checkpoint — not from the start.
Run on your own account — AWS, Azure, GCP, or any provider — with no VaultLayer per-run charge. Your credits and contracts apply.
The run needed more hours than any machine under it lived. It finished anyway — because the run doesn't depend on any one machine surviving.
72B QLoRA fine-tune, one H200-class GPU at a time. The machines under it turned over 10 times — instance recycles, a full disk, a stall — and the job resumed from its last checkpoint every time. Measured, not modeled.
| VaultLayer managed — what it actually charged | $111.56 |
| Same hours at on-demand list price, same GPU class | $217.26 |
The cheap raw capacity underneath isn't an alternative for a 28-hour job — this run's machines turned over 10 times along the way. The system resumed from checkpoint every time. That's what you're paying for.
"The babysitting tax. … I'd lose 20–30 minutes per session to wrapper work around the actual training … The fine-tune finished cleanly — adapter saved, no leftover pod to clean up, no manual sync, no babysitting."
✓ Fails closed: jobs error hard if credentials are missing — no silent fallback
✓ Checkpoint-resume validated end-to-end across our routing pool
✓ Managed checkpoint credentials are job-scoped and short-lived
In our cost model, restarting from scratch compounds with run length until raw spot costs more than on-demand around day three — while checkpoint-and-resume grows roughly linearly, its effective hourly rate staying stable. Assumptions are labeled on the chart; the full model lives in our guides.
Managed runs are billed at the provider's rate plus a flat percentage platform fee shown before the run starts. BYOC is a monthly platform contract — no VaultLayer per-run charges; your cloud bills you directly for compute.
Only what's needed to run your job: the credentials you provide, used to provision and tear down instances. It fails closed if none exist — no silent fallback.
Checkpoints save on your cadence and sync continuously. On interruption, the job restarts from the last checkpoint on fresh hardware. Integration is usually auto-inserted on the first run.
Training and fine-tuning — PyTorch, JAX, and Hugging Face. If it runs with python train.py, it works. Not real-time inference serving.
Sign up, install the CLI with pip install vaultlayer, and submit your first job with vl run train.py. Or email the founder directly — rahuljain@vaultlayer.cloud.
Bring your training job. It finishes even when the hardware doesn't.
Sign upOr email the founder directly — rahuljain@vaultlayer.cloud