Weave Router cuts coding costs 40%
Weave has released Weave Router, a source-available, low-latency model routing proxy for coding agents like Claude Code, Cursor, and Codex. It automatically directs inference requests to the most optimal model based on task complexity, saving up to 40% on token costs.
Weave Router tackles the growing financial burden of frontier models in agentic workflows by replacing manual model toggling with real-time routing. By targeting high-frequency, multi-turn coding agents, it makes long-running development loops economically viable for teams.
- –Uses a local ONNX embedding model to route prompts in under 50ms, avoiding external network latency for routing decisions.
- –Integrates directly as a proxy with support for Anthropic Messages, OpenAI Chat Completions, and Gemini native APIs.
- –Emits OpenTelemetry traces out of the box, allowing developers to visualize decisions and latency in their preferred observability stack.
- –Released as a source-available project under the Elastic License 2.0, providing a self-hosted alternative to closed routing APIs.
- –Addresses the cost surge from tokenizer changes and large context windows in next-generation coding agents.
DISCOVERED
47d ago
2026-06-26
PUBLISHED
47d ago
2026-06-26
RELEVANCE
AUTHOR
adchurch