Dual Blackwell GPUs run 167 GB DeepSeek-V4 FP8
A developer shared a deployment recipe for running the official FP8 version of DeepSeek-V4-Flash-0731 alongside DSpark speculative decoding on a dual NVIDIA RTX PRO 6000 Blackwell (SM120) GPU rig. Requiring approximately 167 GB of VRAM, the model fits cleanly across the system's combined 192 GB VRAM capacity (2× 96 GB) without offloading or truncation.
High-performance local inference for massive open models is becoming practical on workstation hardware thanks to next-generation Blackwell GPUs and FP8 quantization. Fitting a ~167 GB FP8 model natively across dual 96 GB VRAM GPUs eliminates system RAM offloading bottlenecks. Integrating DeepSeek-V4-Flash with DSpark speculative decoding maximizes local throughput and latency performance while providing an actionable hardware blueprint for developers hosting enterprise-grade open LLMs locally.
DISCOVERED
1h ago
2026-08-01
PUBLISHED
1h ago
2026-08-01
RELEVANCE
AUTHOR
light_foundry