It listens to live audio, decides when someone has finished speaking, replies with audio, and returns transcriptions.
PR walkthrough
Your voice model is a live conversation service, not a recording uploaded to your app server.
When a learner speaks, their browser streams small audio packets directly to Gemini Live using a short-lived, tightly constrained ticket. Your server stays in the middle only to approve the session and handle small helper requests such as translations and suggested replies.
It creates reply ideas and readable translations after a turn. It is not the model carrying the spoken conversation.
It reads the microphone, converts audio to the format Gemini expects, plays the response, and animates the mic ring.
Current step
Start practice
1 of 6
Captures speech, shows levels, and plays the tutor.
Checks the user and asks Gemini for a limited session ticket.
Understands speech and generates the spoken response.
The app asks for microphone permission. No live connection exists yet, so no speech has been sent to the model.
A single spoken turn
What happens after you say “I’d like a shawarma”
The browser reads the microphone and measures volume for the pulsing ring.
Audio is reduced to mono 16 kHz PCM and encoded into small messages.
The model detects speech, transcribes it, and chooses an Arabic response.
24 kHz PCM audio returns. The browser queues and plays it immediately.
After a final turn, Gemini Flash can add transliteration and an English gloss.
Responsibilities
Which part is responsible for what?
| Part | What it does | What it does not do |
|---|---|---|
| Browser audio code | Gets mic permission, streams PCM, plays returned audio, manages mute/pause, and drives the visual speech rings. | It does not decide what the tutor says or keep a long-term secret API key. |
| App API | Checks the signed-in user in production, rate-limits access, validates small JSON requests, and obtains a constrained temporary Gemini session. | It is not relaying the whole audio conversation packet by packet. |
| gemini-2.5-flash-native-audio-latest | The live model: real-time listening, turn detection, Arabic tutor voice, audio output, and speech transcriptions. | It is not used for the after-the-fact reply cards or translation cards. |
| gemini-3.6-flash | The helper model: exactly three suggested replies, plus transliteration and English translation for final turns. | It does not receive the continuous microphone stream or generate the tutor’s real-time voice. |
The browser receives permission for one limited live session—not your permanent Gemini secret.
The server creates a short-lived Gemini token constrained to the chosen live model and the session’s audio-only tutor configuration. In production, the request requires a signed-in user. This lets the browser talk to Gemini fast enough for conversation without exposing the main server API key.
The model is deliberately guided.
- It is locked to Palestinian Levantine Arabic, not generic MSA.
- It receives the scenario, tutor role, opening line, phrases, and learner level.
- It has automatic turn detection so the learner can interrupt it naturally.
- It uses the scenario’s selected voice.
Mute and pause are different.
- Mute: stops sending mic audio while keeping the live session ready.
- Pause: ends the active audio stream, stops the rings and tutor playback, and marks the practice as paused.
- Resume: the next mic audio packet reopens the stream.
- Idle: the app pauses the session after 15 seconds without activity.
What this PR fixed
The important changes in plain English
✓ The live connection now fails safely
Any socket error or close—even one with no reason—now stops the microphone, audio player, animation loop, and timers. Old callbacks cannot restart a cleaned-up session.
✓ The mic ring is cheaper to run
Volume state updates are limited to about 20 times a second and only when the volume meaningfully changes, instead of causing a root UI render every animation frame.
✓ Model access is protected
Production requests require sign-in, have per-user rate limits, reject oversized bodies, and return safer errors. Live starts are limited to 6 per 10 minutes; text helpers have tighter minute limits.
✓ Text-model output is structured
Reply suggestions and translations now request JSON with a defined schema. If the model produces invalid JSON, the app returns empty helper content rather than showing raw model text in the UI.