YOU ARE VIEWING ONE ITEM FROM THE AICRIER FEED

Olmo-core 3 opens trillion-parameter MoE training

AICrier tracks AI developer news across Product Hunt, GitHub, Hacker News, YouTube, X, arXiv, and more. This page keeps the article you opened front and center while giving you a path into the live feed.

// WHAT AICRIER DOES

7+

TRACKED FEEDS

24/7

SCRAPED FEED

Short summaries, external links, screenshots, relevance scoring, tags, and featured picks for AI builders.

Olmo-core 3 opens trillion-parameter MoE training
OPEN LINK ↗
// 1h agoOPENSOURCE RELEASE

Olmo-core 3 opens trillion-parameter MoE training

Ai2 released Olmo-core 3, an open training stack for scaling mixture-of-experts language models toward trillion-parameter sizes. Its redesigned distributed architecture delivers up to 2.7× higher throughput in preliminary benchmarks while keeping expert-routing costs manageable.

// ANALYSIS

Olmo-core 3 makes open MoE experimentation meaningfully more credible by publishing the systems work usually hidden inside frontier labs. The headline numbers are promising, but developers should treat them as infrastructure benchmarks—not evidence of model quality.

  • –Expert, pipeline, and distributed-optimizer parallelism reduce per-GPU memory pressure at extreme scale
  • –Ai2 reports 52,000 tokens per second per GPU for a 47B MoE, versus 19,400 with its earlier FSDP implementation
  • –MXFP8 increased throughput about 21% over BF16 while reducing peak active memory in controlled tests
  • –The stack has been benchmarked on a 1.2T-parameter configuration across 512 GPUs, with random routing rather than a trained model
  • –Open access to routing experiments, ablations, and training infrastructure could give academic labs a stronger alternative to proprietary stacks
// TAGS
olmo-core-3moellmtraining-infragpuopen-source

DISCOVERED

1h ago

2026-10-04

PUBLISHED

2h ago

2026-10-04

RELEVANCE

9/ 10

AUTHOR

AI Search