Introducing Gemini 3.5 Transcribe for Advanced Speech-to-Text
Today, we are unveiling Gemini 3.5 Transcribe, our most advanced speech-to-text model designed for intelligent voice interaction. This new model excels where traditional speech recognition systems often falter, particularly with background noise, complex jargon, and removing speech disfluencies. Gemini 3.5 Transcribe transforms raw audio into precise, polished, and formatted text.
The benefits of this transcription model are already evident in our Gemini app and on Android, offering enhanced voice capabilities such as Rambler. Now, developers can integrate these features using Gemini 3.5 Transcribe through the Gemini API available in Google AI Studio and the Gemini Enterprise Agent Platform.
Seamless Integration for Developers
Gemini 3.5 Transcribe is designed to integrate smoothly into developer workflows, whether the focus is on building voice agents, real-time captioning, or post-call analytics. The model is accessible via two distinct APIs, catering to various development needs.
This model captures natural speech patterns to better understand user intent and recognize specialized vocabulary. It is capable of handling live language switches and offers seamless streaming transcription. Additionally, it cleans up speech disfluencies efficiently.
Advanced Capabilities
Gemini 3.5 Transcribe provides transcription with multi-speaker recognition and word-level timestamps. It represents a significant improvement over our previous model, Chirp 3, enhancing capabilities, reducing word error rates, and offering better latency. According to Artificial Analysis, the time to finalize transcription has improved by 70%. On the FLEURS benchmark, the model delivers precise multilingual performance, improving upon Chirp 3, with a word error rate of 5.50% in streaming mode and 5.04% in non-streaming scenarios.
Besides integration with the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, Gemini 3.5 Transcribe enhances everyday experiences across Google platforms like Gboard, Antigravity, the Gemini app, and Chrome by capturing nuances, intentions, and inline edits effortlessly.
The model allows users to analyze files, generate images, and conduct searches in the Gemini app on macOS using voice commands. It also utilizes the Rambler feature on Android to automatically remove filler words and clean up speech.
Utilizing the Gemini Live API, developer platforms such as Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents can create and deploy high-performance voice-driven interfaces efficiently. These platforms handle complex real-time media streaming infrastructure, allowing developers to focus on perfecting the user experience.
Companies like Vivo, Intellitek Health, and Lingopal have praised Gemini 3.5 Transcribe for its impressive latency, accuracy, and wide-ranging language support.
Sign up for our newsletters to receive updates on products, events, special offers, and more. Check your inbox to confirm your subscription, or subscribe using a different email address. Your information will be used according to Google’s privacy policy, and you may opt out at any time.
