Qwen3.8-Flash Goes Free on B.AI
B.AI is offering Alibaba’s hosted Qwen3.8-Flash API at 0 Credits, with multimodal input, agent controls, and a 1M-token context window. Chat access is rolling out separately, and the free offer is temporary.
Free access makes Qwen3.8-Flash an unusually compelling model for prototyping long-context agents, though developers should avoid designing around a promotion that may disappear.
- –The 1M-token context window can handle large repositories, document sets, screenshots, and video in fewer retrieval steps.
- –Function calling, structured output, caching, and hosted tools make it practical for production-style agent experiments.
- –The related Qwen3.8-Flash-Next release exposes a 125B-parameter MoE architecture with only 6B active per token, but the hosted Flash endpoint’s parameters and benchmark results are not independently published. Qwen GitHub repository: https://github.com/QwenLM/Qwen3.8-Flash-Next
- –Developers should benchmark latency, rate limits, and reliability before migrating workloads from established paid APIs.
DISCOVERED
55m ago
2026-08-27
PUBLISHED
1h ago
2026-08-27
RELEVANCE
AUTHOR
thenameissky1