Google Unveils Gemini 3.5 Transcribe Speech Model
The announcement positions Gemini 3.5 Transcribe against a competitive field where its non-streaming error rate of 2.6% ranks fifth on Artificial Analysis's leaderboard, behind models like ElevenLabs' Scribe v2 and Microsoft's MAI-Transcribe-1.5, making it a benchmark comparison rather than a clear top performer.
Key Facts
- Google announced Gemini 3.5 Transcribe on August 26, 2026 (early August 27 Japan time).
- Gemini 3.5 Transcribe achieves a word error rate of 4.0% for streaming and 2.6% for non-streaming, as measured by Artificial Analysis.
- The model supports over 85 languages and allows custom vocabulary with up to 1,000 custom terms.
- Pricing is estimated at about $0.005 per minute for recorded audio and $0.009 per minute for real-time transcription, with free tiers.
- Gemini 3.5 Transcribe is already available in Gboard's AI Voice Input on Pixel 11 series devices and the macOS Gemini app's Speak to Window feature.
Reporting from 2 sources: GameBusiness.jp, GIGAZINE.
Google announced Gemini 3.5 Transcribe, a speech recognition model that converts raw audio into clean, formatted text in real time while automatically removing filler words like "um" and "uh." The model also reflects speaker self-corrections and applies formatting such as bullet points and parentheses, turning rambling speech into polished prose. It is available through two APIs: the Live API for continuous bidirectional streaming with sub-second latency, and the Interactions API for recorded audio with speaker identification and word-level timestamps. Google reports a word error rate of 4.0% for streaming and 2.6% for non-streaming, as measured by Artificial Analysis, and a 70% reduction in time to final transcription compared to the earlier Chirp 3 model. The model supports over 85 languages and allows custom vocabulary for technical terms. Gemini 3.5 Transcribe is already integrated into Android's Rambler voice input and the macOS Gemini app's Speak to Window feature. Developers can access it through Google AI Studio, Google Antigravity, and the Gemini Enterprise Agent Platform. Pricing is estimated at about $0.005 per minute for recorded audio and $0.009 per minute for real-time, with free tiers available.
- GIGAZINE frames the model for voice agents, real-time captioning tools, and call analysis pipelines, and notes it "represents a significant advancement" over Chirp 3 in Artificial Analysis measurements.
- GameBusiness.jp details a side-by-side test where the conventional voice input misheard phone names as "Google Pixelmator" and "iPhone 7 Pro Max," while Gemini 3.5 Transcribe read context correctly.
- The same test shows the model keeps a filler word when it is contextually necessary, rendering "いわゆるフィラー、あーとかうーとかが" as "いわゆるフィラー(あー、うー)が" rather than deleting it.
- Speaker separation supports up to 3 speakers reliably, with 4 or more experimental; the API documentation allows up to 8 speakers.
- Audio length is capped at 1 hour, reduced to 30 minutes when speaker separation or timestamps are enabled.
- Developers can register up to 1,000 custom vocabulary terms.
- Beyond transcription, the model supports function calling, such as having another model generate images based on content.
- Gboard's AI Voice Input (Rambler) is a key feature of the Pixel 11 series and also supports editing specific places by voice or rewriting the style of entire text after input.
- macOS Gemini app access uses the Fn/globe key as a shortcut by default, working in text input areas of any app or window.
- At the May 2026 Gemini Intelligence announcement, AI voice input was cited as a feature with rollout promised to the latest Galaxy devices.
- Chrome and Gemini Enterprise for Customer Experience are listed as "coming soon" with timing undecided.
Synthesized by Yomimono from the 2 cited sources below, including Japanese-language reporting where cited, then editorially reviewed before publishing.