In the present day, we’re introducing Gemini 3.5 Transcribe, our most exact speech-to-text mannequin but, designed for clever voice interactions. Not like standard speech recognition fashions that wrestle with background noise, complicated jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts uncooked audio straight into correct, polished, formatted textual content.
Throughout our merchandise just like the Gemini app and on Android, we’ve seen customers already benefiting from this transcription mannequin with new voice capabilities like Rambler on Android and within the Gemini app on macOS. Now, builders can construct comparable capabilities with Gemini 3.5 Transcribe within the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.
We have constructed 3.5 Transcribe to plug seamlessly into your developer workflows, whether or not you’re constructing voice brokers, real-time captioning instruments, or post-call analytics pipelines. The mannequin is out there throughout two separate APIs:
Actual-time streaming: Delivers steady, bidirectional streaming with sub-second latency for interactive voice apps by way of the Dwell API utilizing gemini-3.5-transcribe-live.Pre-recorded audio processing: Transcribes recorded audio, conferences, name logs, and extra with speaker attribution and word-level timestamps by way of the Interactions API utilizing gemini-3.5-transcribe.
Get extra exact and clever transcription
Gemini 3.5 Transcribe is designed to seize your pure talking fashion to raised perceive your intent and acknowledge customized vocabulary, so you possibly can execute duties together with your voice.
Good transcription: Seamlessly handles self-corrections (like “let’s meet Tuesday—no, Wednesday”), removes filler phrases (“ums” and ‘“ahs”), auto-formats your textual content.Perform calling: The mannequin can delegate complicated duties (resembling picture era and file evaluation) to different Gemini fashions by way of perform calls. Presently accessible within the Gemini macOS app.Extra exact transcription: As measured by Synthetic Evaluation, achieves a mean Phrase Error Price (WER) of 4.0% for streaming and a couple of.6% for non-streaming use-cases. It exhibits sturdy efficiency throughout noisy, real-world environments, precisely capturing alphanumeric entities like postal codes and order IDs.Customized vocabulary: Acknowledges specialised jargon and distinctive spellings by seamlessly adapting transcriptions to your offered customized vocabulary.World language help: Mechanically detects and transcribes over 85 languages, seamlessly dealing with regional accents and various dialects.Multi-speaker identification: Precisely attributes speech in pre-recorded audio with timestamps for as much as three audio system (help for 3+ audio system is experimental).

