Nemotron 3.5 Lightning Hits Merge Gateway
NVIDIA’s 30B Mixture-of-Experts model, with just 3B active parameters, is now available for free through Merge Gateway. It offers a 1M-token context window and is positioned for high-throughput agent workloads.
Nemotron 3.5 Lightning makes open-weight inference more practical by optimizing for sustained speed rather than sheer parameter count.
- –Sparse activation should reduce serving costs and latency for high-volume agent tasks
- –A 1M-token context window enables long-running workflows and large RAG workloads
- –Merge Gateway availability removes deployment friction for developers testing the model
- –NVIDIA’s claimed 4x output-rate advantage and 86% PinchBench score are promising, but independent evaluations will matter
- –The model’s value will depend on tool calling, reliability, and quality under concurrency—not just benchmark throughput
DISCOVERED
3h ago
2026-08-12
PUBLISHED
23h ago
2026-08-11
RELEVANCE
AUTHOR
merge_api