weights.zip · Model weight compression

Your models, one-third smaller. Every bit still perfect.

Lossless compression built exclusively for AI model weights: 34% smaller files that restore byte-for-byte identical, decompress at GPU speed, and never, ever expand.

Restores byte-for-byte identical Decodes faster than NVMe
Install the CLI ·
curl -fsSL https://weights.zip/install.sh | bash
Request early access →
What weights.zip does

Lossless compression, built only for model weights.

weights.zip shrinks neural-network checkpoints by up to 34 percent, and by up to 87 percent on pruned or sparse models. Every compressed file restores bit-for-bit identical to the original, verified on the way back out, so what you load is exactly what you trained.

At scale the math is hard to ignore. Store 100 PB of checkpoints and you keep roughly 66 PB, trimming about 500,000 dollars a month off a typical hot-storage bill. The savings compound across every replica, every snapshot, and every version you retain.

Decompression runs at GPU speed, outrunning the NVMe drive the weights sit on, so loads get faster instead of slower. Because compressed weights stay resident in GPU memory, a single card can serve a model roughly a third larger than its VRAM would otherwise allow.

It drops into the workflow you already have. Point it at your safetensors files, turn on verify mode, and ship. In head-to-head tests it beat every alternative on ratio while decoding faster than all of them. Smaller bills. Faster loads. Zero loss.

Two ways to run weights.zip

Compress in place, or let us host the registry.

In your own stack

CLI and Python library

Point the CLI or the library at your safetensors and checkpoints. Compression and byte-exact restore run inside your own infrastructure. Nothing leaves your network.

Get the SDK →
Managed registry

Hosted weight registry

Push versioned, compressed checkpoints to a private per-account registry and pull them back at GPU speed. Resident-in-VRAM serving and verify-on-restore are built in.

See the registry →
Typical compression

How much smaller, by model type. Always lossless.

Dense
Full-precision fp16 and bf16 checkpoints.
34%smaller, lossless
Restores byte-for-byte identical.
Quantized
int8 and fp8 export weights.
30%smaller, lossless
Exact bits, faster loads.
Pruned / sparse
Magnitude-pruned and structured-sparse.
87%smaller, lossless
Packs down furthest of all.
Enterprise
Multi-model registries, resident-in-VRAM serving, custom SLA.
Customquoted on your bill
Priced against your storage.
Savings calculator

See what you would save on weight storage, in seconds.

$ / mo
5.0
Every byte restores identical. Verified on decompress.
Current / yr
$2,400,000
With weights.zip / yr
$1,584,000
You save / yr
$816,000 34% lower
Over 3 years
$2.45M
Over 10 years
$8.16M
Footprint after
3.3 PB

At this setting weights.zip does not reduce your bill. If your weights already live on the cheapest cold tier, savings may be small.

Byte-for-byte restore
Every checkpoint decompresses bit-for-bit identical to the original, verified on restore.
GPU-speed decode
Decompression outruns your NVMe drive, so weights load faster, not slower.
Never expands
Output is always smaller than input, or the file passes through untouched. It never, ever grows.
Drops into your workflow
Native safetensors support and a verify mode. Point it at your files and ship.
Request early access

Figures are estimates anchored to your current storage bill. Your final savings are confirmed by a no-cost pilot that compresses a sample of your own checkpoints and verifies a byte-exact restore. Savings depend on your current storage class; if your weights already sit on the cheapest cold tier you may see little.