Updated: 2026-08-28
Two Minute Papers recently covered DeepSeek's V4 Pro 0813 release, and the headline is blunt: an open-weight model is now trading blows with closed frontier systems. This post unpacks what the release actually changes, how close it really is, and — the practical part — what kind of hardware can run a model of this class at home.
What is DeepSeek V4 Pro 0813?
V4 Pro 0813 is DeepSeek's August snapshot of its V4-generation model family, and two design choices stand out. First, the production checkpoints ship quantization-aware, so the official weights are already compact rather than shrunk as an afterthought. Second, DeepSeek publishes a hybrid attention design that keeps long-context serving tractable. Like its siblings, it is a Mixture-of-Experts model: every expert sits in memory but only a few activate per token, which shifts the bottleneck from compute to storage.
How close is it to closed models?
Benchmark chatter around 0813 focuses on agentic and coding work, where DeepSeek's vendor-reported results put it within striking distance of closed frontier systems. Independent replication is still catching up, so treat specific scores as provisional. The trajectory is not provisional, though: our earlier DeepSeek V4 benchmark breakdown showed the previous snapshot closing gaps that looked permanent months ago. The cost angle matters just as much — we also covered how DeepSeek tackled AI's billion-dollar inference problem with architectural efficiency rather than brute force.
Can your PC run it?
Here is the sobering part. Because a Mixture-of-Experts model must store every expert, combined RAM plus VRAM decides what fits: quantized community builds span roughly 100 GB at heavy compression up to around 170 GB for lossless weights, according to Compare AI Hardware's DeepSeek V4 hardware requirements breakdown. No single consumer GeForce card holds the full model even at heavy quantization. Self-hosting therefore lives on high-unified-memory machines or multi-GPU workstations, with partial CPU offload as the slow-but-cheap fallback.
VRAM and memory options compared
| Hardware | Memory | Bandwidth | Role for V4-class self-hosting |
|---|---|---|---|
| Mac Studio (M3 Ultra) | up to 512 GB unified | 819 GB/s | Simplest single-box option |
| RTX 6000 Ada | 48 GB GDDR6 | 960 GB/s | Multi-GPU workstation build |
| RTX A6000 | 48 GB GDDR6 | 768 GB/s | Prior-generation workstation |
| H100 | 80 GB HBM3 | 3350 GB/s | Datacenter inference workhorse |
| H200 | 141 GB HBM3e | 4800 GB/s | Highest-capacity tracked option |
The table mirrors the hardware records Compare AI Hardware tracks for the V4 family. A high-memory desktop without a discrete GPU can run heavily quantized builds on the processor alone — the small active-parameter count keeps that viable — but token speed will test your patience.
Frequently asked questions
Is DeepSeek V4 Pro 0813 free to use?
The weights are open, so downloading and self-hosting costs nothing in licensing fees. Your real costs are hardware, power, and setup time; hosted API access is priced separately by providers.
How much VRAM does DeepSeek V4 Pro need?
Think combined memory, not just VRAM: roughly 100 GB for heavily quantized builds and up to about 170 GB for lossless weights, per the hardware breakdown linked above.
Can a single RTX 4090 run it?
No. A 24 GB card cannot hold the stored experts even at heavy quantization; it can only participate in multi-GPU or CPU-offload setups.
What is DSpark?
DSpark is the speculative-decoding technique demonstrated in the episode linked below — it pairs a small draft model with the large one to raise token throughput.
V4 Pro or V4 Flash — which should I run?
Flash is the lighter, faster sibling and far easier to host; Pro trades convenience for capability. Match the tier to your memory budget before deciding.
Sources and further reading
- DeepSeek V4 Pro 0813 (official): https://ift.tt/LYU6jXT
- Episode sponsor — Lambda GPU Cloud: https://ift.tt/Z8SFAru
- DSpark full episode: https://www.youtube.com/watch?v=1yBU41auQhw
- Original source links: https://ift.tt/iP6QqZG, https://ift.tt/uBXoN5Y, https://ift.tt/lxnZRcN, https://ift.tt/WhpIRbk, https://ift.tt/pgaH8Yq, https://ift.tt/B9jHd2X, https://ift.tt/XvLqtRm
🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible:
Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan Stankye, Shawn Becker, Steef, Taras Bobrovytsky, Tazaur Sagenclaw, Tybie Fitzhugh, Ueli Gallizzi
Disclosure: machine-learning.null.pictures and compareaihardware.com are operated by the same team. Links to compareaihardware.com are editorial recommendations, not paid placements.
No comments:
Post a Comment