Facts
- Release
- game · 2026-08-28 · 2026-08-28
- Noted
- It is available through two APIs: the Live API for continuous bidirectional streaming with sub-second latency, and the Interactions API for recorded audio with speaker identification and word-level timestamps. · 2026-08-28
- Noted
- Google reports a word error rate of 4.0% for streaming and 2.6% for non-streaming, as measured by Artificial Analysis, and a 70% reduction in time to final transcription compared to the earlier Chirp 3 model. · 2026-08-28
- Noted
- The model supports over 85 languages and allows custom vocabulary for technical terms. · 2026-08-28
- Noted
- Gemini 3.5 Transcribe is already integrated into Android's Rambler voice input and the macOS Gemini app's Speak to Window feature. · 2026-08-28
- Noted
- Developers can access it through Google AI Studio, Google Antigravity, and the Gemini Enterprise Agent Platform. · 2026-08-28
- Noted
- Pricing is estimated at about $0.005 per minute for recorded audio and $0.009 per minute for real-time, with free tiers available. · 2026-08-28
- Noted
- The announcement positions Gemini 3.5 Transcribe against a competitive field where its non-streaming error rate of 2.6% ranks fifth on Artificial Analysis's leaderboard, behind models like ElevenLabs' Scribe v2 and Microsoft's MAI-Transcribe-1.5, making it a benchmark comparison rather than a clear top performer. · 2026-08-28
Structured graph also available as JSON at /public/entities/gemini-3-5-transcribe.
CC BY 4.0.
1h ago
Google announced Gemini 3.5 Transcribe, a speech recognition model that converts raw audio into clean, formatted text in real time while automatically removing filler words like "um" and "uh." The model also reflects speaker self-corrections and applies formatting such as bullet points and parentheses, turning rambling speech into polished prose. It is available through two APIs: the Live API for continuous bidirectional streaming with sub-second latency, and the Interactions API for recorded audio with speaker identification and word-level timestamps. Google reports a word error rate of 4.0% for streaming and 2.6% for non-streaming, as measured by Artificial Analysis, and a 70% reduction in time to final transcription compared to the earlier Chirp 3 model. The model supports over 85 languages and allows custom vocabulary for technical terms. Gemini 3.5 Transcribe is already integrated into Android's Rambler voice input and the macOS Gemini app's Speak to Window feature. Developers can access it through Google AI Studio, Google Antigravity, and the Gemini Enterprise Agent Platform. Pricing is estimated at about $0.005 per minute for recorded audio and $0.009 per minute for real-time, with free tiers available.