JevBench maps typed-decision model tradeoffs.
JevBench v1 evaluates Jev-class decision models across accuracy, cost, latency, calibration, reliability, and openness. Its 19 September run tested nine systems on 242 typed decisions each, with GPT-5.6 Luna and Jev 1.13.0 leading accuracy. [Benchmark Heaven](https://benchmarkheaven.com/jev-models/v1)
JevBench’s strongest contribution is treating decision models as production components, where speed, confidence, cost, and reproducibility matter alongside raw accuracy.
- –GPT-5.6 Luna reached 97.1% accuracy, while Jev 1.13.0 scored 96.3%; overlapping confidence intervals make this a close comparison rather than a definitive podium.
- –Reporting five separate axes avoids hiding meaningful tradeoffs behind one composite score.
- –The benchmark measures typed outputs and probabilities, making calibration and integration reliability more relevant than prose-generation quality.
- –Its 242-item mix includes published, held-out, and imported decisions, but the small system pool and single-server setup limit how broadly results should be generalized.
- –The MIT-licensed harness and published decisions give developers a practical starting point for reproducing or extending the evaluation.
DISCOVERED
1h ago
2026-09-27
PUBLISHED
1h ago
2026-09-27
RELEVANCE
AUTHOR
DIY Smart Code