Moonshot AI releases 2.8T Kimi-K3 open weights
Moonshot AI has published open weights for Kimi-K3, a 2.8-trillion parameter sparse Mixture-of-Experts transformer model activating 16 of 896 experts per token. Built for complex reasoning and long-horizon tasks, it features native vision support, a 1-million-token context window, and innovations like Kimi Delta Attention and Attention Residuals.
The open release of Kimi-K3 represents a major shift in the frontier AI landscape by bringing trillion-parameter MoE architecture into the open-weights ecosystem.
- –**Massive MoE Efficiency**: Uses extreme sparsity with 16 active experts out of 896, keeping per-token compute manageable despite a huge 2.8T parameter total.
- –**Long-Context Multimodality**: Combines a 1M token context window with native vision capabilities, targeting complex agentic and document-heavy workflows.
- –**Architectural Breakthroughs**: Introduces Kimi Delta Attention and Attention Residuals to improve training stability and scaling efficiency at massive parameter dimensions.
- –**Open-Source Disruption**: Places strong competitive pressure on closed commercial model providers by enabling developers and enterprises to run frontier-grade intelligence independently.
DISCOVERED
1h ago
2026-07-27
PUBLISHED
3h ago
2026-07-27
RELEVANCE
AUTHOR
nateb2022