GPT-5.6 Sol–Luna Routing Stretches Codex Quota
A developer reports running Codex for eight hours with GPT-5.6 Sol at Extra High and Luna Max handling subagents, using just 29% of the weekly allowance. The result suggests quota-aware model routing can make long-running agent workflows more economical, though it remains an anecdotal measurement.
The bigger story is specialization: Sol handles orchestration and judgment while Luna absorbs bounded execution. That could materially increase completed work per quota, but usage varies by plan, model, effort level, caching, and task complexity.
- –OpenAI positions Sol as its flagship model and Luna as its fastest, most affordable tier, with selectable effort levels in Codex. [OpenAI overview](https://openai.com/index/gpt-5-6/)
- –A practical split is Sol for architecture and review, Luna for parallel exploration, implementation, and routine fixes.
- –Codex usage depends on task size, complexity, model, and execution surface—not simply hours spent. [OpenAI Help Center](https://help.openai.com/en/articles/11369540-getting-started-with-codex)
- –The 29% figure is a useful field signal, not a guaranteed quota multiplier; allowances and accounting can differ across plans.
- –Developers should log model, effort, token/cache mix, request count, and subagent count before changing their defaults.
DISCOVERED
1h ago
2026-08-30
PUBLISHED
1h ago
2026-08-30
RELEVANCE
AUTHOR
rudrank