← all stories

Gemini 3 5 Transcribe

Facts

Release
game · 2026-08-28 · 2026-08-28
Noted
It is available through two APIs: the Live API for continuous bidirectional streaming with sub-second latency, and the Interactions API for recorded audio with speaker identification and word-level timestamps. · 2026-08-28
Noted
Google reports a word error rate of 4.0% for streaming and 2.6% for non-streaming, as measured by Artificial Analysis, and a 70% reduction in time to final transcription compared to the earlier Chirp 3 model. · 2026-08-28
Noted
The model supports over 85 languages and allows custom vocabulary for technical terms. · 2026-08-28
Noted
Gemini 3.5 Transcribe is already integrated into Android's Rambler voice input and the macOS Gemini app's Speak to Window feature. · 2026-08-28
Noted
Developers can access it through Google AI Studio, Google Antigravity, and the Gemini Enterprise Agent Platform. · 2026-08-28
Noted
Pricing is estimated at about $0.005 per minute for recorded audio and $0.009 per minute for real-time, with free tiers available. · 2026-08-28
Noted
The announcement positions Gemini 3.5 Transcribe against a competitive field where its non-streaming error rate of 2.6% ranks fifth on Artificial Analysis's leaderboard, behind models like ElevenLabs' Scribe v2 and Microsoft's MAI-Transcribe-1.5, making it a benchmark comparison rather than a clear top performer. · 2026-08-28

Structured graph also available as JSON at /public/entities/gemini-3-5-transcribe. CC BY 4.0.

All coverage

1h ago

Google Unveils Gemini 3.5 Transcribe Speech Model

Google announced Gemini 3.5 Transcribe, a speech recognition model that converts raw audio into clean, formatted text in real time while automatically removing filler words like "um" and "uh." The model also reflects speaker self-corrections and applies formatting such as bullet points and parentheses, turning rambling speech into polished prose. It is available through two APIs: the Live API for continuous bidirectional streaming with sub-second latency, and the Interactions API for recorded audio with speaker identification and word-level timestamps. Google reports a word error rate of 4.0% for streaming and 2.6% for non-streaming, as measured by Artificial Analysis, and a 70% reduction in time to final transcription compared to the earlier Chirp 3 model. The model supports over 85 languages and allows custom vocabulary for technical terms. Gemini 3.5 Transcribe is already integrated into Android's Rambler voice input and the macOS Gemini app's Speak to Window feature. Developers can access it through Google AI Studio, Google Antigravity, and the Gemini Enterprise Agent Platform. Pricing is estimated at about $0.005 per minute for recorded audio and $0.009 per minute for real-time, with free tiers available.