Google launches Gemini 3.5 Transcribe
Google’s new speech-to-text model converts raw audio into polished, formatted text, handling filler words, self-corrections, custom vocabulary, and 85+ languages. It is available through Google AI Studio, the Gemini API, and Gemini Enterprise Agent Platform.
Gemini 3.5 Transcribe’s real advantage is intent-aware transcription rather than raw word accuracy, positioning speech recognition as a developer-ready layer for voice agents and productivity apps.
- –Live API streaming targets interactive voice experiences with sub-second latency.
- –Interactions API supports recorded audio, speaker attribution, and word-level timestamps.
- –Custom vocabulary and dialect support should make it more practical for specialized domains.
- –Artificial Analysis measures word error rates of 4.0% for streaming and 2.6% for non-streaming use cases, though developers should benchmark their own audio.
- –Broad API availability could pressure dedicated transcription providers, especially for teams already using Gemini.
DISCOVERED
1h ago
2026-08-26
PUBLISHED
2h ago
2026-08-26
RELEVANCE
AUTHOR
JackWoth98