Jev 1.13 beats Cohere, Voyage in latency
Michael Chomsky shared latency benchmark results comparing TypeSafe AI's Jev 1.13 model against leading dedicated rerankers for search and retrieval pipelines. In the test, Jev 1.13 registered the fastest speeds with a 149 ms median and 205 ms p95 latency, beating Cohere v3.5 (176 ms median / 357 ms p95), Voyage 2.5 Lite (183 ms median / 232 ms p95), and Cohere 4 Fast (183 ms median / 824 ms p95).
Fast structured-decision models are threatening to disrupt the dedicated cross-encoder reranking market sooner than expected.
- –Jev 1.13 edges out established reranking offerings on both median speed (149 ms) and tail latency (205 ms p95), which is crucial for real-time RAG stacks where latency budgets are tight.
- –Jev demonstrates much tighter tail variance compared to Cohere 4 Fast, which experienced severe tail latency spikes reaching 824 ms at p95.
- –Because Jev avoids token-by-token generation overhead and provides calibrated probability scores, it can natively perform binary relevance checks and document filtering at high concurrency without the cost penalty of standard LLMs.
- –While latency numbers are very promising, comprehensive quality benchmarks evaluating nDCG@10 and MRR on datasets like BEIR are still needed to verify whether its ranking quality matches or exceeds top specialized cross-encoders.
DISCOVERED
1h ago
2026-09-21
PUBLISHED
1h ago
2026-09-21
RELEVANCE
AUTHOR
michael_chomsky