Nimble lands in Ollama with typed decisions
Bespoke Labs’ Apache 2.0 Nimble is a 9B Qwen3.5-9B fine-tune for fast local classification, routing, moderation, and scoring. Ollama now exposes it through the /v1/systemone API, where it reports 75.7% accuracy across 13 public datasets—just behind Jev’s 76.0%.
Nimble makes a compelling case for replacing unnecessary text generation with bounded, probability-backed decisions, though its benchmark parity with Jev is highly task-dependent.
- –Local inference reduces latency, cost, and data exposure for routing and moderation workloads.
- –Directly scoring allowed answer tokens produces typed outputs without parsing generated JSON.
- –Nimble leads Jev on some tasks, but trails sharply on moderation and German-language routing.
- –Its 9B footprint and flat-schema limitations make it less flexible than a general-purpose LLM.
- –The reported results are promising, but developers should benchmark their own policies before automating high-impact decisions.
DISCOVERED
1h ago
2026-09-30
PUBLISHED
1h ago
2026-09-30
RELEVANCE
AUTHOR
DIY Smart Code