Two separate APIs—Live for real-time voice agents (gemini-3.5-transcribe-live), Interactions for post-recorded audio with speaker attribution—both handle custom vocabulary and 85+ languages natively.
Summary
Replaces Chirp 3 with 70% latency reduction and measurable accuracy gains (4.0% WER streaming, 2.6% non-streaming). Enables voice-first workflows in existing Gemini API integrations without architectural changes.
Why it matters
Replaces Chirp 3 with 70% latency reduction and measurable accuracy gains (4.0% WER streaming, 2.6% non-streaming). Enables voice-first workflows in existing Gemini API integrations without architectural changes.
Implementation verdict
Public preview now in Google AI Studio and Gemini API. Requires switching from Chirp 3 model identifiers and choosing between Live (streaming) or Interactions (batch) paths. Ready to try for voice agents and transcription pipelines; production readiness depends on your WER tolerance and language mix.
Sources
Dev Signal
Get briefs like this in your inbox — free, every weekday.
100+ sources compressed into one 4-minute read. Ranked, cited, implementation-ready.