Cekura Bench Ranks Voice Models on Real Calls
Cekura Bench tests realtime voice models as complete phone agents across 82 healthcare scenarios, repeating each three times to measure reliability, accuracy, stalled calls, latency, and cost. Public transcripts, tool calls, and benchmark code make the results independently inspectable.
Cekura’s strongest contribution is treating voice evaluation as an end-to-end systems problem rather than a leaderboard of isolated model capabilities.
- –GPT Realtime 2.1 leads reliability at 79.3%, while Phonic v1 delivers the fastest median response at 1.60 seconds.
- –The cascade baseline reaches 82.9% reliability, beating every native speech-to-speech model and challenging the assumption that direct audio models are automatically superior.
- –Repeated-run scoring exposes consistency failures that one-off demos hide, especially on longer Medicare intake calls.
- –Public transcripts let developers investigate why models fail instead of trusting opaque aggregate scores.
- –The benchmark gives teams a practical framework for choosing models based on quality, speed, and cost tradeoffs.
DISCOVERED
2h ago
2026-10-08
PUBLISHED
8h ago
2026-10-08
RELEVANCE
AUTHOR
[REDACTED]