NVIDIA Groq 3 LPX hits production
NVIDIA’s Groq 3 LPX inference accelerator has entered full production, with Nebius integrating it into Token Factory for faster agentic AI generation. SpaceXAI also plans to deploy NVIDIA Vera CPUs for orchestration, tool use, code execution, and simulation.
NVIDIA is turning inference into a heterogeneous systems game: GPUs handle context, LPUs handle decode, and CPUs keep agent loops moving. The strategic message is clear—agent performance now depends as much on token latency and orchestration as raw accelerator throughput.
- –LPX combines 256 Groq 3 LPUs with Vera Rubin GPUs, targeting low-latency decoding for long-context and multi-step agent workloads. [NVIDIA technical blog](https://developer.nvidia.com/blog/inside-nvidia-groq-3-lpx-the-low-latency-inference-accelerator-for-the-nvidia-vera-rubin-platform)
- –NVIDIA reports 3,400 output tokens per second on Gemma 4 31B with a 100,000-token context, though the result is vendor-sponsored and workload-specific. [NVIDIA](https://blogs.nvidia.com/blog/vera-rubin-lpx-spectrum-x-nvlink-fusion/)
- –Nebius’ Token Factory gives developers a production API path to this hardware, making LPX relevant beyond hyperscaler-only deployments.
- –Vera CPUs acknowledge that agents spend substantial compute outside the model call: coordinating tools, executing code, moving data, and running simulations.
- –The Groq integration shows NVIDIA defending its inference moat by absorbing specialized silicon into a tightly controlled full-stack platform.
DISCOVERED
1d ago
2026-08-25
PUBLISHED
1d ago
2026-08-24
RELEVANCE
AUTHOR
AI_AgentFL