AI inference costs plunge 47% per quarter
Alex Tabarrok highlights an Epoch AI report by Luke Emberson and David Roodman showing that AI inference costs for a fixed capability level have declined 47% per quarter over the past three years. This 13-fold annual deflation outpaces historical technologies like compute and DNA sequencing, with benchmark costs on GPQA Diamond falling over 700-fold in under 18 months.
Hyper-deflation in AI inference costs threatens to upend current capital expenditure assumptions—if raw intelligence rapidly approaches zero marginal cost, massive data center moats may provide diminishing returns far faster than investors anticipate.
- –**Unprecedented Deflationary Velocity:** A 13-fold annual reduction in inference cost dwarfs Moore's Law, fundamentally shifting the bottleneck from hardware availability to application-layer integration.
- –**Rapid Reasoning Commoditization:** Benchmarks such as GPQA Diamond demonstrate that frontier-grade reasoning becomes micro-penny commodity compute in under two years.
- –**Evolving Open vs. Closed Dynamics:** Proprietary labs are competing aggressively on inference efficiency rather than raw capability alone, narrowing the operational cost window that open models traditionally exploited.
- –**Edge Viability Approaching:** As compute efficiency accelerates, models matching current frontier capabilities will feasibly run on standard consumer hardware within three to five years, dampening long-term cloud dependency.
DISCOVERED
1h ago
2026-09-23
PUBLISHED
4h ago
2026-09-23
RELEVANCE
AUTHOR
gotmedium