Cerebras Serves Qwen 3.8 27B at 1,500 Tokens/s
Cerebras now offers Qwen3.8-27B on public inference endpoints at approximately 1,500 tokens per second, with 64K free-tier and 128K paid-tier context limits. The 27B open-weight vision-language model supports image/video understanding and controllable reasoning.
This deployment makes a capable dense model feel dramatically more practical for interactive agent workflows, though token efficiency still matters more than headline throughput.
- –1,500 tokens per second could make coding agents and research loops feel nearly instantaneous.
- –Cerebras’ hosted context limits are below Qwen3.8-27B’s native 262K-token capability, so developers should verify deployment-specific constraints.
- –Thinking is enabled by default; tuning reasoning effort or disabling it for simple tasks will reduce unnecessary latency and cost.
- –The speed is notable because it approaches Cerebras’ larger Qwen offering, but GPT OSS 120B remains listed at roughly 3,000 tokens per second.
- –Community testing suggests the model can overthink and consume long traces, making workload-specific time-to-solution benchmarks essential. [Hacker News discussion](https://news.ycombinator.com/item?id=49299605)
DISCOVERED
1h ago
2026-09-03
PUBLISHED
3h ago
2026-09-03
RELEVANCE
AUTHOR
altertable