GLM-5.3-Flash Goes Open, Cuts Inference Costs
Z.ai has open-sourced GLM-5.3-Flash, its first natively multimodal GLM-5 model, with 320B total parameters and 18B active parameters. Released under MIT licensing, it claims GLM-5.2-beating performance at one-tenth the price.
GLM-5.3-Flash makes efficient, open multimodal models far more credible for production workloads, though its enormous total parameter count still complicates local deployment.
- –Hybrid sparse and linear attention targets lower long-context serving costs
- –Native image and video support expands the model beyond coding-only use cases
- –Z.ai reports performance approaching Claude Opus 4.8 on coding and agent benchmarks
- –Open weights, MIT licensing, vLLM, and SGLang support make experimentation straightforward
- –The one-tenth pricing claim is compelling, but most benchmark and cost comparisons remain company-reported
DISCOVERED
4h ago
2026-08-26
PUBLISHED
6h ago
2026-08-26
RELEVANCE
AUTHOR
Philpax
