Brood War Bench pits frontier LLMs in StarCraft
Brood War Bench is an AI evaluation project by Ben Swerdlow testing large language models in real-time StarCraft: Brood War matches across 171 games. Codex Astra topped the leaderboard undefeated with disruptive probe tactics, while results highlighted that real-time environments severely penalize reasoning latency as lower-effort configurations often outperformed higher-effort ones.
Real-time agency exposes the heavy latency penalty of reasoning models, proving that in fast-moving environments, operational frequency beats deep deliberation.
- –Reasoning latency is fatal: Models like Grok 4.6 generated thousands of reasoning tokens while issuing minimal command batches, often getting wiped out before fielding a combat unit.
- –Disruption beats macro: Top performer Codex Astra won largely through basic disruption and worker harassment rather than sustained macro production or advanced tactical coordination.
- –Lower effort often outperforms: Lower and medium reasoning effort configurations frequently beat maximum effort settings because higher APM and timely reactive actions prevented models from freezing.
- –Subagent fragmentation: When models spawned subagents for economy and military control, lack of inter-agent coordination led to classic beginner mistakes like trickling units in one by one.
DISCOVERED
1h ago
2026-09-20
PUBLISHED
1h ago
2026-09-20
RELEVANCE
AUTHOR
steipete