Modal Ships Moonshot Kimi K3 on Merge Gateway
Modal has launched support for Moonshot AI's Kimi K3 on its Merge Gateway platform, integrated with Modal's custom DFlash speculative decoding architecture. The deployment reaches up to 460 tokens per second, offering a 360% increase in interactive speed and an 88% throughput boost on agentic workloads.
Custom speculative decoding stacks are fast becoming the key performance differentiator for cloud infrastructure providers serving real-time AI agents.
* Ultra-Fast Inference: Achieving 460 tokens per second significantly reduces latency, enabling near-instantaneous feedback in conversational and interactive AI interfaces.
* Tailored for Agentic Workloads: An 88% increase in throughput directly addresses performance bottlenecks common in multi-turn, multi-step agent execution pipelines.
* Infrastructure Optimization: Proprietary hardware and speculative decoding tweaks allow Modal to extract maximum efficiency from model deployment without relying solely on raw compute scaling.
DISCOVERED
2h ago
2026-07-28
PUBLISHED
2h ago
2026-07-28
RELEVANCE
AUTHOR
merge_api