NCP-ArchPreview cuts pretraining tokens, outperforms OLMo-3-7B
NCP-ArchPreview is an open 8.9-billion-parameter latent-space language model that predicts multi-token discrete concepts and feeds them back into token-level generation rather than relying purely on next-token prediction. Matching OLMo-3-7B's final pretraining loss using only 51.3% of training tokens and finishing 2.45 points higher downstream, the release includes open weights, recipes, and checkpoints.
Next Concept Prediction directly attacks the compute inefficiency of token-by-token autoregression, potentially disrupting pretraining economics if scaling laws hold at larger capacities. Achieving comparable pretraining loss with ~51% of training tokens represents a massive reduction in the compute budget required to train competitive models, while downstream gains like +5.99 on GSM8K indicate that higher-level conceptual structures enhance reasoning. Open-sourcing weights, checkpoints, and recipes allows the community to validate stability and test concept bottlenecks, though open questions remain around inference latency, codebook collapse risks, and multimodal adaptation.
DISCOVERED
1h ago
2026-09-13
PUBLISHED
1h ago
2026-09-13
RELEVANCE
AUTHOR
mark_k
