B.AI Makes GLM-5.3-Flash Free
Z.AI’s GLM-5.3-Flash is a 320B-parameter, 18B-active mixture-of-experts model with native text, image, video, and file understanding, a 1M-token context window, and MIT-licensed weights. B.AI is temporarily offering its API at zero credits, making a frontier-class multimodal model unusually accessible for developer experimentation.
The economics are the real story: sparse activation, efficient attention, and aggressive distribution make this look less like a conventional model launch and more like an attempt to reset expectations for inference pricing.
- –Only 18B parameters activate per token, helping explain how a 320B model can target “Flash” economics while retaining frontier-level ambitions.
- –Hybrid sparse and linear attention reportedly cuts attention compute and KV-cache requirements, directly attacking the cost of long-context agent workloads.
- –Native multimodality makes it useful for visual software engineering, document analysis, UI inspection, and agents that need to verify rendered output.
- –The MIT weights expand deployment options, but the model still requires roughly hundreds of gigabytes of accelerator memory, keeping serious self-hosting out of reach for most developers.
- –B.AI’s free access is best viewed as a high-leverage trial and ecosystem-acquisition strategy; developers should expect pricing and availability to change after the promotion.
DISCOVERED
1h ago
2026-08-29
PUBLISHED
2h ago
2026-08-29
RELEVANCE
AUTHOR
kidsreallycute