PR walkthrough

Your voice model is a live conversation service, not a recording uploaded to your app server.

When a learner speaks, their browser streams small audio packets directly to Gemini Live using a short-lived, tightly constrained ticket. Your server stays in the middle only to approve the session and handle small helper requests such as translations and suggested replies.

The live model Gemini Live is the tutor’s ears and voice.

It listens to live audio, decides when someone has finished speaking, replies with audio, and returns transcriptions.

The helper model Gemini Flash handles small text jobs.

It creates reply ideas and readable translations after a turn. It is not the model carrying the spoken conversation.

The browser Your app does the audio plumbing.

It reads the microphone, converts audio to the format Gemini expects, plays the response, and animates the mic ring.

Current step

Start practice

1 of 6

BrowserMicrophone + audio player

Captures speech, shows levels, and plays the tutor.

Your serverApproval gate

Checks the user and asks Gemini for a limited session ticket.

Gemini LiveLive Arabic tutor

Understands speech and generates the spoken response.

The browser starts the session.

The app asks for microphone permission. No live connection exists yet, so no speech has been sent to the model.

A single spoken turn

What happens after you say “I’d like a shawarma”

1. Mic capture

The browser reads the microphone and measures volume for the pulsing ring.

2. PCM conversion

Audio is reduced to mono 16 kHz PCM and encoded into small messages.

3. Gemini Live

The model detects speech, transcribes it, and chooses an Arabic response.

4. Tutor audio

24 kHz PCM audio returns. The browser queues and plays it immediately.

5. Transcript helper

After a final turn, Gemini Flash can add transliteration and an English gloss.

Responsibilities

Which part is responsible for what?

PartWhat it doesWhat it does not do
Browser audio codeGets mic permission, streams PCM, plays returned audio, manages mute/pause, and drives the visual speech rings.It does not decide what the tutor says or keep a long-term secret API key.
App APIChecks the signed-in user in production, rate-limits access, validates small JSON requests, and obtains a constrained temporary Gemini session.It is not relaying the whole audio conversation packet by packet.
gemini-2.5-flash-native-audio-latestThe live model: real-time listening, turn detection, Arabic tutor voice, audio output, and speech transcriptions.It is not used for the after-the-fact reply cards or translation cards.
gemini-3.6-flashThe helper model: exactly three suggested replies, plus transliteration and English translation for final turns.It does not receive the continuous microphone stream or generate the tutor’s real-time voice.
Why the ticket matters

The browser receives permission for one limited live session—not your permanent Gemini secret.

The server creates a short-lived Gemini token constrained to the chosen live model and the session’s audio-only tutor configuration. In production, the request requires a signed-in user. This lets the browser talk to Gemini fast enough for conversation without exposing the main server API key.

Tutor behaviour

The model is deliberately guided.

  • It is locked to Palestinian Levantine Arabic, not generic MSA.
  • It receives the scenario, tutor role, opening line, phrases, and learner level.
  • It has automatic turn detection so the learner can interrupt it naturally.
  • It uses the scenario’s selected voice.
Controls

Mute and pause are different.

  • Mute: stops sending mic audio while keeping the live session ready.
  • Pause: ends the active audio stream, stops the rings and tutor playback, and marks the practice as paused.
  • Resume: the next mic audio packet reopens the stream.
  • Idle: the app pauses the session after 15 seconds without activity.

What this PR fixed

The important changes in plain English

✓ The live connection now fails safely

Any socket error or close—even one with no reason—now stops the microphone, audio player, animation loop, and timers. Old callbacks cannot restart a cleaned-up session.

✓ The mic ring is cheaper to run

Volume state updates are limited to about 20 times a second and only when the volume meaningfully changes, instead of causing a root UI render every animation frame.

✓ Model access is protected

Production requests require sign-in, have per-user rate limits, reject oversized bodies, and return safer errors. Live starts are limited to 6 per 10 minutes; text helpers have tighter minute limits.

✓ Text-model output is structured

Reply suggestions and translations now request JSON with a defined schema. If the model produces invalid JSON, the app returns empty helper content rather than showing raw model text in the UI.