Gemini 4 Pro leaks reveal 256K output limit
Google's next-generation Gemini 4 Pro model has surfaced through internal leaks under the codename "Argon," displaying an expansive 256K output token capacity alongside an expected 2M context window. Early footage from internal testing showcases generation times reaching approximately 2.4 minutes under high thinking effort, highlighting Google's aggressive push toward compute-intensive, long-horizon test-time reasoning and massive multi-file output generation.
A 256K output token ceiling signals a decisive industry transition from conversational AI to autonomous generation of entire software repositories and exhaustive documentation in a single pass. A 2.4-minute generation time indicates deep iterative thinking loops, positioning Argon directly against leading frontier reasoning models. Expanding maximum completions to 256K tokens effectively unlocks monolithic codebase synthesis and deep research artifact production without synthetic chunking. Combining long context with high output windows demonstrates that Google DeepMind's infrastructure optimizations are keeping pace with soaring inference demands.
DISCOVERED
1h ago
2026-09-15
PUBLISHED
1h ago
2026-09-15
RELEVANCE
AUTHOR
WorldofAI