Olmo-core 3 opens trillion-parameter MoE training
Ai2 released Olmo-core 3, an open training stack for scaling mixture-of-experts language models toward trillion-parameter sizes. Its redesigned distributed architecture delivers up to 2.7× higher throughput in preliminary benchmarks while keeping expert-routing costs manageable.
Olmo-core 3 makes open MoE experimentation meaningfully more credible by publishing the systems work usually hidden inside frontier labs. The headline numbers are promising, but developers should treat them as infrastructure benchmarks—not evidence of model quality.
- –Expert, pipeline, and distributed-optimizer parallelism reduce per-GPU memory pressure at extreme scale
- –Ai2 reports 52,000 tokens per second per GPU for a 47B MoE, versus 19,400 with its earlier FSDP implementation
- –MXFP8 increased throughput about 21% over BF16 while reducing peak active memory in controlled tests
- –The stack has been benchmarked on a 1.2T-parameter configuration across 512 GPUs, with random routing rather than a trained model
- –Open access to routing experiments, ablations, and training infrastructure could give academic labs a stronger alternative to proprietary stacks
DISCOVERED
1h ago
2026-10-04
PUBLISHED
2h ago
2026-10-04
RELEVANCE
AUTHOR
AI Search