General Legal unveils real-world legal benchmark
AI-native law firm General Legal has unveiled a benchmark evaluating whether frontier AI models can produce client-ready legal work across real client matters rather than merely fulfilling isolated checklist criteria. Testing 13 frontier models against actual attorney deliverables, the study revealed that while most models cleared 80% of individual rubric items, nearly half achieved an all-pass rate of just 10% without lawyer intervention.
Achieving 90% compliance on an isolated legal rubric is practically worthless when the remaining 10% still requires a supervising attorney to review and rewrite the entire matter from scratch.
• Pass rates mask critical failures: Most frontier models comfortably clear 80% of individual checklist criteria, but nearly half achieved an all-pass rate of just 10%, underscoring the divide between modular accuracy and deliverable viability.
• Reasoning scales only capable baselines: Extended reasoning effort yielded massive performance gains for capable models like GPT-6 Astra (jumping from 2 passes at Low to 7 at Ultra), whereas weaker models like GPT-5.6 Terra burned tokens without passing additional matters.
• Higher inference cost does not ensure quality: Pricing ranged from $0.07 to $6 per matter, with expensive models like Fable 5.1 and Claude Opus 5 underperforming significantly cheaper configurations.
• Holistic judgment over checklist execution: Modern legal evaluations must move beyond synthetic Q&A to track strategic trade-offs, extraneous additions, and negotiation dynamics in full context.
DISCOVERED
1h ago
2026-09-24
PUBLISHED
1h ago
2026-09-24
RELEVANCE
AUTHOR
gen_legal_inc