Splash powers fast local Mac agents
Splash is an open-source Apple Silicon inference engine that connects local models to Codex, OpenCode, Claude Code, and compatible APIs. It uses model-specific Metal kernels, DFlash 2 speculative decoding, caching, and batching to accelerate on-device coding workloads. [GitHub](https://github.com/incoai/splash)
Splash makes a strong case for specialized local inference, though its Apple Silicon, macOS, memory, and supported-model requirements limit its audience.
- –Benchmarks report 2–5× decode gains over llama.cpp on supported Qwen models, with larger benefits under concurrent agent workloads. [GitHub releases](https://github.com/incoai/splash/releases)
- –OpenAI- and Anthropic-compatible APIs make migration straightforward for existing coding-agent workflows.
- –Model-specific optimization trades universal flexibility for substantially better performance and simpler tuning.
- –M3-or-newer hardware and at least 36 GB of unified memory make this a premium local-development tool.
- –The project’s Apache-2.0 license and integration with LM Studio broaden its potential beyond a niche CLI. [Inco AI](https://inco.ai/blog/splash/)
DISCOVERED
1h ago
2026-10-01
PUBLISHED
1h ago
2026-10-01
RELEVANCE
AUTHOR
Github Awesome