Google launches Gemini 3.5 Flash-Lite
Google launched Gemini 3.5 Flash-Lite, its fastest and most cost-effective 3.5 model delivering speeds up to 350 output tokens per second. Optimized for low-latency agentic search and bulk document processing, it is available via Gemini API and Google AI Studio starting at $0.30 per million input tokens.
Google is aggressively targeting the high-throughput developer market by undercutting competitors on speed and pricing to position Gemini as the default runtime for autonomous agents.
- –At 350 output tokens/second, this model directly addresses the compounding latency issue inherent in multi-step AI agent workflows.
- –The highly competitive pricing ($0.30/1M input, $2.50/1M output) makes large-scale processing tasks like bulk document parsing and RAG highly viable.
- –Integrating the model directly into Google Search highlights a strategic shift toward utilizing faster, specialized model variants to power real-time consumer search features.
DISCOVERED
3h ago
2026-07-21
PUBLISHED
4h ago
2026-07-21
RELEVANCE
AUTHOR
GoogleAIStudio