Geo News
Community curated by people like you
LatestAICryptoHealthWorld AffairsUS Politics
Google releases Gemini 3.5 Transcribe speech-to-text model with 2.6% error rate across 85+ languages
00

Google releases Gemini 3.5 Transcribe speech-to-text model with 2.6% error rate across 85+ languages

Aug 26, 2026

Google has released Gemini 3.5 Transcribe, an AI speech-to-text model that automatically filters out filler words and corrects speech stumbles in real time. The model achieves a 2.6% word error rate on pre-recorded audio and a 4.0% rate on live streams, delivering transcriptions 70% faster than its predecessor, Chirp 3. It supports over 85 languages and is available via two distinct developer endpoints, powering consumer features like Gboard's Rambler on the Pixel 11 and the macOS Gemini app.

Gemini 3.5 Transcribe release

  • ▪Google released Gemini 3.5 Transcribe alongside two companion models, Gemini 3.5 Live and Gemini 3.5 Live Experimental, which together form the "Gemini Audio" brand.
  • ▪Google released Gemini 3.5 Transcribe, an AI-powered speech-to-text model designed to streamline voice input by editing out disfluencies and formatting text on the fly.

Speech cleaning capabilities

  • ▪Gemini 3.5 Transcribe features a "smart" mode that automatically removes disfluencies like "ums" and "ahs," resolves spoken self-corrections inline, and applies structured formatting.
  • ▪Gemini 3.5 Transcribe features a "verbatim" mode that returns all spoken elements, including fillers, repetitions, and false starts.
  • ▪Gemini 3.5 Transcribe allows users to edit transcribed text using voice commands and adapt the output to a custom vocabulary list of up to 1,000 terms.

Word error rate metrics

  • ▪On the multilingual FLEURS benchmark, Gemini 3.5 Transcribe reports a word error rate of 5.50% for streaming and 5.04% for non-streaming audio, compared to 7.32% for the previous Chirp 3 model on live speech.
  • ▪Google reported that Gemini 3.5 Transcribe achieves an average word error rate of 4.0% on streaming audio and 2.6% on pre-recorded files, as measured by Artificial Analysis.
  • ▪Gemini 3.5 Transcribe improves the time to final transcription by 70% compared to Google's previous transcription engine, Chirp 3.

Two-endpoint API architecture

  • ▪The gemini-3.5-transcribe-live endpoint caps continuous streaming sessions at 10 minutes and does not support speaker diarization or word-level timestamps.
  • ▪The gemini-3.5-transcribe endpoint supports speaker diarization for up to three speakers and word-level timestamps, but limits audio files to 30 minutes when these features are enabled.
  • ▪Google deployed Gemini 3.5 Transcribe as two distinct API endpoints: gemini-3.5-transcribe-live for sub-second streaming and gemini-3.5-transcribe for pre-recorded files.

Application deployment options

  • ▪Gemini 3.5 Transcribe is available to developers in public preview through Google AI Studio and Google Antigravity, and to enterprises via the Gemini Enterprise Agent Platform.
  • ▪Google plans to expand Gemini 3.5 Transcribe support to Chrome voice typing, Search Live, Gemini Live, Docs, Keep, and Gmail.
  • ▪Gemini 3.5 Transcribe powers the "Rambler" dictation feature on Gboard for the Pixel 11 and is integrated into the Gemini app on macOS.

Language support coverage

  • ▪Gemini 3.5 Transcribe supports automatic language detection and regional accents across more than 85 languages, including mid-sentence code-switching.
  • ▪Gemini 3.5 Transcribe's language coverage stretches past 85 locales, which Google notes is critical for markets where speakers routinely switch languages mid-sentence.

4 sources

Timesofindia
Google's Gemini 3.5 Transcribe AI model edits your speech as you talk, and Chrome gets it next
View source article
Digitaltrends
Gemini 3.5 Transcribe wants you to stop worrying about being a perfect talker
View source article
Marktechpost
Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Languages
View source article
Arstechnica
Google announces Gemini 3.5 Transcribe for AI-powered speech-to-text
View source article

Featured stories

View more in Speech recognition

Linux kernel vulnerabilities surge to nearly 2,000 per release as AI bug hunters overwhelm maintainers

Sep 1, 2026 · 1 source

Dyson launches $499 AI-powered toothbrush with built-in camera

Sep 1, 2026 · 3 sources

OpenAI advertising business reaches $1 billion annual run rate

Aug 31, 2026 · 2 sources

Anthropic unveils MHS standard for AI agents to operate physical devices

Aug 27, 2026 · 5 sources

Story comments

Loading comments…

Related Projects

Google

Topics

Speech recognitionMultimodal modelsAI tools & productsAI voice & speech tools

Featured stories

View more in Speech recognition

Linux kernel vulnerabilities surge to nearly 2,000 per release as AI bug hunters overwhelm maintainers

Sep 1, 2026 · 1 source

Dyson launches $499 AI-powered toothbrush with built-in camera

Sep 1, 2026 · 3 sources

OpenAI advertising business reaches $1 billion annual run rate

Aug 31, 2026 · 2 sources

Anthropic unveils MHS standard for AI agents to operate physical devices

Aug 27, 2026 · 5 sources