Gemini Live Avatar Configuration

Authentication
Avatar setup
50
Voice Activity Detection (VAD) & Turn-Taking
Wait time before model starts replying
Audio buffer prepended to speech
Context
Tools

ℹ️ Disabled by default. When enabled, tools demonstrate async function calling and scheduling.

Header

⚠️ Avoid mid-session edits (impacts prefix caching & latency). Only use for rare, high-priority static information (e.g. storage approvals).

System Instructions / Persona

Available variables: $name (Avatar name), $mood (Mood), $voiceDescription (Voice style), $language (Language), $location (Coordinates), $userGivenName (User first name), $userFamilyName (User last name), $otherAvatarName (Other avatar name)

ℹ️ System instructions are non-editable during an active session. Any persona or prompt context that may need updating mid-session should go into the Footer.

Footer

ℹ️ Use for all dynamic information from the system prompt that can change mid-session.

Custom Avatar configuration

Hint: You can create a custom avatar by taking a photo, uploading an image, or generating one from a prompt. For best results with camera input, ensure a uniform background.

Create Custom Voice
⚠️
Sensitive Technology Consent Required Voice cloning is a highly sensitive technology. You MUST obtain explicit, documented consent from the person whose voice is being cloned before recording and using their voice sample.
Select script to read during recording:

"When the sunlight strikes raindrops in the air, they act like a prism and form a rainbow. The rainbow is a division of white light into many beautiful colors. These take the shape of a long round arch."

🎙️ Tips for the best 15-second recording:
  • Speak naturally: Don't use your "radio voice" or over-enunciate. Read it exactly as you would speak to a friend.
  • Background noise: Limit background noise as much as possible. Turn off fans, AC units, and close windows.
  • Smile slightly!: Smile during the audio recording to add warmth to the replicated voice.
Step 1: Record Voice Sample
Not recorded
Step 2: Record Consent Statement
Speak this sentence exactly: "I am the owner of this voice and I consent to Google using this voice to create a synthetic voice model."
Not recorded
Advanced Settings
Visible Controls:
Multiples of 128 recommended. Lower = less latency, Higher = less overhead.
Session Statistics

General

RTT: - ms
CPU Pressure: -

Avatar 1

Initial latency:
Total: - ms
Latency:
Avg: - ms (Curr: - ms)
Session Duration: - s
Transparency: - ms (avg: - ms, -, -)
Video Playback: 0 dropped frames
Packet Jitter: - ms
Video Codec: Auto
Interruptions: 0
Message Metric Sent Received
Audio Messages 0 0
Video Messages 0 0
Text Messages 0 0
Total Messages 0 0
Token Metric Tokens
Prompt Tokens (Input) 0
Model Tokens (Output) 0
Total Tokens 0
⌨️

Keyboard Control Active

Press Ctrl + S (or Cmd ⌘ + S on Mac) to start and stop the session.

Click or press any key to dismiss