Codex runs Dwarf Fortress autonomously for 29 hours
Developer @WhySoLivid executed a 29-hour long-horizon AI agency field study by orchestrating Dwarf Fortress through OpenAI's Codex, with token consumption totaling approximately 1.9 billion reported tokens. While the automated fortress survived the continuous run, the experiment underscored severe friction in human monitoring and safety bureaucracy, illustrating how governance frameworks can become a key operational backpressure point during sustained autonomous AI execution.
Testing AI agents in open-ended environment simulations like Dwarf Fortress shows that safety governance and human oversight become the main operational bottleneck during extended autonomous runs.
- –Proves long-horizon agent execution capability across a massive 1.9B token scale and 29-hour runtime.
- –Illustrates how traditional human safety monitoring struggles to keep pace with continuous high-throughput AI agent actions.
- –Highlights the necessity of building scalable governance tools that handle monitoring backpressure without stalling autonomous agent performance.
DISCOVERED
2h ago
2026-07-27
PUBLISHED
7h ago
2026-07-26
RELEVANCE
AUTHOR
WhySoLivid