Turn live and recorded speech into precise, formatted text
Gemini 3.5 Transcribe is Google's speech-to-text model for intelligent voice applications. Developers can stream audio with sub-second latency or process recordings with speaker attribution and word-level timestamps for voice agents, captions, meetings, call analysis, and other transcription workflows.
Route voice AI sessions to models measured by language and cost
Deploy durable AI agents with a real Linux computer per session
Data is being prepared
Traffic data is refreshed monthly.
Run a persistent AI agent that remembers, learns, and acts across tools.
Extend your IDE with development agents from Google Antigravity.