Qwen3.8-2.4T Raises Open-Weight Bar
Alibaba’s Qwen3.8-2.4T-A95B brings 2.4 trillion total parameters, 95 billion active per token, and up to 1.01 million tokens of context to an open-weight model. It is the first Qwen Max-tier model available for developer deployment.
Qwen is making frontier-scale capability accessible while MoE sparsity keeps inference somewhat practical—but the hardware requirements still limit true self-hosting to well-funded teams.
- –512 experts with 11 activated per token deliver a roughly 25:1 sparsity ratio
- –Native 262K context, extendable to about 1.01M tokens, suits long-running coding and research agents
- –Day-one vLLM support and FP4 quantizations improve deployment flexibility across NVIDIA and AMD hardware
- –The model requires multiple datacenter-class GPUs, so “open weights” does not mean consumer-friendly
- –Open access to a Max-tier checkpoint could accelerate distillation, fine-tuning, and independent evaluation
DISCOVERED
2h ago
2026-08-22
PUBLISHED
6h ago
2026-08-22
RELEVANCE
AUTHOR
honozcom