GPT-6.1 Sol debuts third on MacroscopeBench
Macroscope’s code-review benchmark places GPT-6.1 Sol at #3, behind GPT-6 Sol on recall but with 30% fewer comments and high precision. It also slightly beats GPT-6 Astra at roughly one-seventh the cost.
GPT-6.1 Sol’s advantage is signal per dollar, not absolute bug coverage. Fewer, more accurate comments could make it a stronger default for high-volume automated review.
- –MacroscopeBench measures real-bug recall, precision, signal-to-noise, cost, and review duration.
- –Lower recall means teams may still need stronger models for critical repositories.
- –Reduced comment volume can improve reviewer trust and reduce alert fatigue.
- –GPT-6.1 Sol costs $2/$10 per million input/output tokens, compared with GPT-6 Astra’s $10/$50 pricing.
- –Benchmark-specific results should be validated against each team’s own pull-request history.
DISCOVERED
1h ago
2026-10-02
PUBLISHED
1h ago
2026-10-02
RELEVANCE
AUTHOR
Macroscope