日本語 ← Back to home
AI Tools

Google Releases Gemini 3.8 Live and 3.5 Transcribe for Voice Applications

Google's new Gemini audio models support live speech interactions and transcription in more than 85 languages.

Article ID: TC-0033 Published:

On September 15, 2026, Google introduced new audio models through the Gemini API and Google AI Studio: Gemini 3.8 Live, Gemini 3.8 Live Extended Thinking and Gemini 3.5 Transcribe.

Gemini 3.8 Live is designed for speech-to-speech interaction while carrying out tasks. This supports voice-first applications where dialogue and actions can occur within the same experience.

DIFFERENT MODELS FOR DIFFERENT JOBS: Google's announcement separates conversational speech models from dedicated transcription. Gemini 3.8 Live focuses on spoken interaction, while Gemini 3.5 Transcribe converts audio into text. Developers should choose according to the task rather than assume one model is best for every audio workflow.

SPEECH-TO-SPEECH INTERACTION: A voice-first application must preserve the rhythm of conversation. A direct speech-oriented design can differ from a pipeline that transcribes input, processes text and synthesizes a reply. However, speech-to-speech architecture does not guarantee low latency in every network or workload.

INTERRUPTIONS AND TURN-TAKING: Natural conversations include interruptions, corrections and changes of direction. Voice assistants need to recognize when a user is speaking and decide whether to pause or revise a response. Interrupting too easily can be as frustrating as continuing to speak over the user.

EXTENDED THINKING: Gemini 3.8 Live Extended Thinking is positioned for more complex requests. Additional reasoning may be useful when comparing several conditions or planning multiple steps. Developers still need to balance response quality against conversational speed. The right configuration depends on user expectations and the consequences of an incorrect answer.

TRANSCRIPTION AS A DISTINCT TASK: Gemini 3.5 Transcribe is described as supporting more than 85 languages. Potential applications include meeting records, captions and voice input. Language coverage does not establish equal accuracy for specialized vocabulary, regional accents or every acoustic environment.

MULTIPLE SPEAKERS AND PROPER NOUNS: Meetings can contain overlapping speech, company names and technical terminology. These conditions can make transcription difficult. For important records, users need a way to check uncertain passages against the original recording rather than relying on an unreviewed transcript.

MEASURING THE USER EXPERIENCE: Useful metrics include the delay between a completed utterance and the start of a response, how well the system handles interruptions and how much effort is needed to correct mistakes. A strong benchmark score does not automatically produce a natural conversation.

NETWORK CONDITIONS MATTER: Real-time audio depends on connectivity. Variable latency or temporary disconnections can disrupt the flow of dialogue. Applications intended for mobile use should test reconnection, recovery and the handling of partially transmitted speech, not only performance on a stable office network.

AUDIO PRIVACY: Recordings may contain names, private discussions and background information about the user's environment. Applications should make recording visible and explain retention, deletion and transmission practices. Workplace meeting tools may also need to account for participant consent and organizational policies.

DESIGNING FOR CORRECTION: Recognized text should not always be treated as final. Mistakes involving names, numbers and instructions can have practical consequences. Interfaces that let users inspect and edit transcripts, or revisit the source audio, can make the overall system more dependable.

CHOOSING AN IMPLEMENTATION: Teams should first decide whether their main goal is conversation or transcription. They can then compare latency, supported languages, noisy-audio performance, cost and data handling. Pilot testing should use representative recordings and realistic interaction patterns.

THE BIGGER OPPORTUNITY: More capable audio models may bring conversational control into applications traditionally organized around screens and menus. Adoption will depend on more than model intelligence: turn-taking, corrections, consent and recovery from connection problems will shape the experience.

Extended Thinking targets requests requiring more complex reasoning. Google cites strong benchmark results, but a leaderboard position is not a guarantee of performance in every deployed application.

Gemini 3.5 Transcribe is a dedicated speech-to-text model supporting more than 85 languages. Potential uses include meeting notes, captions and voice interfaces.

Developers should evaluate latency, background noise, transcription errors, user consent and handling of sensitive audio rather than focusing on model accuracy alone.

Source

Google Developers Blog (September 15, 2026) ↗