Google Launches Gemini 3.5 Transcribe to Automate Polished Speech-to-Text
The new model moves beyond verbatim transcription by removing filler words and self-corrections in real-time.
Google has introduced Gemini 3.5 Transcribe, a high-precision speech-to-text model designed for both real-time and pre-recorded audio. The release marks a strategic shift toward "intelligent" transcription that prioritizes polished output over literal dictation.
The model is engineered to automatically remove filler words, handle speaker self-corrections, and auto-format text. According to data from Artificial Analysis, the model achieves an average Word Error Rate (WER) of 4.0% for streaming use-cases and 2.6% for non-streaming applications. On the FLEURS benchmark, it recorded a 5.50% WER in streaming mode and 5.04% in non-streaming mode. Performance gains are significant; Google reports that the time to final transcription has improved by 70% compared to the previous Chirp 3 model.
Integration and Accessibility
Gemini 3.5 Transcribe is now available to developers through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. To support different technical needs, Google offers two distinct API paths: 'gemini-3.5-transcribe-live' for sub-second latency streaming and 'gemini-3.5-transcribe' for pre-recorded audio. The latter includes speaker attribution for up to eight speakers, although support for more than three speakers remains experimental.
The technology is already being integrated into the Gemini macOS app and Gboard (Rambler), with a rollout to Chrome expected soon. The model is highly versatile in its linguistic reach, supporting the automatic detection and transcription of more than 85 languages.
The Shift to Smart Transcription
This launch replaces the Chirp 3 model and is part of a broader effort to embed multimodal AI across Google's ecosystem. By moving away from verbatim transcription, Google is attempting to solve the "disfluency problem"—the tendency of human speech to be cluttered with "ums," "ahs," and mid-sentence pivots that make raw transcripts difficult to read.
By automating the cleanup of these errors, Google is positioning voice as a viable primary input for professional drafting and complex workflows. Rather than treating voice-to-text as a convenience for short queries or simple notes, the focus is now on creating a tool capable of producing a clean, professional first draft directly from spoken word.
Future Outlook
As Gemini 3.5 Transcribe integrates further into the Chrome browser and mobile interfaces, the industry will be watching how it handles complex, multi-speaker environments in real-world settings. While the core transcription capabilities are confirmed, the long-term goal is a deeper integration with other Gemini models to enable advanced function calling and context-aware processing based on audio input.