Vercel adds GLM 5.3 FlashX to AI Gateway
Vercel announced that Z.ai's GLM 5.3 FlashX is now available on Vercel AI Gateway, offering high-speed serving at roughly 200 tokens per second. The model retains the multimodal vision processing, tool-calling proficiency, and 1M token context window of GLM 5.3 Flash while reducing generation latency for interactive agent workflows.
Throughput and latency have become just as critical as raw benchmark scores for agentic workflows where multi-step tool loops demand immediate feedback.
- –**Accelerating Agent Loops:** Serving at ~200 TPS substantially compresses round-trip times in tool-heavy coding agents, making complex refactors and reasoning cycles feel responsive rather than sluggish.
- –**No Compromise on Scale:** Combining high throughput with a 1M token context window allows agents to digest entire codebases and visual assets without truncation or throughput bottlenecks.
- –**Zero-Friction Adoption:** With Vercel's zero-markup model and automated CLI setup for popular developer agents, switching to high-speed alternatives from proprietary frontier models is practically frictionless.
DISCOVERED
1h ago
2026-09-18
PUBLISHED
1h ago
2026-09-18
RELEVANCE
AUTHOR
vercel_dev
