Google Gemini 3.5 Transcribe removes filler words from audio
The new AI model automatically detects 85+ languages, formats text, and adapts to specialized vocabulary while the promised Gemini 3.5 Pro remains delayed.
Google has launched Gemini 3.5 Transcribe, a new AI model that automatically removes filler words like "um" and "uh" from audio transcriptions while detecting more than 85 languages and specialized terminology.
The transcription model represents what Google describes as a major advancement over its previous Chirp 3 model, particularly in multilingual performance and reducing wording error rates. The Verge first reported the launch details.
Core capabilities
Gemini 3.5 Transcribe allows users to provide customized vocabulary lists, enabling the model to automatically adapt to unique spelling requirements and industry-specific jargon. This feature aims to reduce manual editing by teaching the AI to recognize specialized terms that might otherwise be transcribed incorrectly.
The model also automatically formats text and can attribute speech to up to three speakers in pre-recorded audio. Word-level timestamps provide precise tracking of when specific words appear in recordings.
According to Google, the system enables users to "edit naturally with just your voice," suggesting voice-driven editing workflows beyond basic transcription.
Availability and platform support
Gemini 3.5 Transcribe began rolling out today in English for macOS Gemini app users and through the Rambler dictation feature on Android in select countries and languages. Developers can access the model in public preview through the Gemini API via AI Studio and Antigravity.
Google indicated that Chrome support will arrive soon, though no specific timeline was provided.
Why it matters
Accurate transcription with automatic filler word removal addresses a persistent pain point for professionals who record meetings, interviews, or content. The ability to customize vocabulary for specialized fields—from medical terminology to technical jargon—could significantly reduce post-production editing time. However, the launch arrives as Google continues to delay its Gemini 3.5 Pro model, originally promised for June release, raising questions about the company's AI rollout timeline and prioritization.
The missing Gemini 3.5 Pro
The transcription model launch follows the earlier release of Gemini 3.5 Live Translate and comes while users continue waiting for Gemini 3.5 Pro, which Google had committed to releasing in June. The company has not provided an updated timeline for that more advanced model.
Google initially mentioned two additional Gemini Audio models—3.5 Live and 3.5 Live Experimental—in pre-publication materials but later clarified that only 3.5 Transcribe was being announced.
Details of the launch were first reported by The Verge.
This is an original analysis by the Omega editorial team. Source reporting: The Verge.
Want systems like this working for your business?
Book a Call