NVIDIA Vera Rubin Hits 10x Blackwell Efficiency
Early measured silicon benchmark results shared by CoreWeave show that NVIDIA's upcoming Vera Rubin architecture provides a massive generational efficiency gain over the GB200 Blackwell NVL72. Running the DeepSeek R1 model, Vera Rubin demonstrated up to 10x higher token throughput per megawatt while preserving equivalent per-user latency and responsiveness.
Vera Rubin's massive efficiency leap shifts the primary constraint on scaling large reasoning models from power availability back to compute deployment.
- –Delivers up to 10x more tokens per megawatt compared to Blackwell NVL72 when running DeepSeek R1.
- –Maintains per-user responsiveness, proving efficiency gains do not require compromising interactive latency.
- –Crucial milestone for cloud infrastructure providers like CoreWeave contending with tight data center power allocations.
DISCOVERED
2h ago
2026-07-22
PUBLISHED
3h ago
2026-07-22
RELEVANCE
AUTHOR
mark_k
