Gemini 3.8 Live can now answer as a talking video avatar, and Google bills the video only while it speaks
Google made Gemini 3.8 Live with Live Avatar generally available in Gemini Enterprise on September 24, 2026, with US and EU endpoints. The model replies through a lip-synced video character while it listens, watches a camera or a shared screen, and calls tools in the background. Avatar video costs $1 per million tokens at 6,192 tokens for each second of speech. By our arithmetic, that comes to about $0.37 per spoken minute before audio, input, and context charges.
Source
Introducing Gemini 3.8 Live with Live AvatarGoogle: The Keyword, AI · Original published September 24, 2026
Google has turned its live dialogue model into a video agent. Gemini 3.8 Live, released on September 15, 2026, now comes with Live Avatar, a character that is generated as it talks, with lips and expressions that follow the speech. Google Cloud says the feature is generally available in Gemini Enterprise and through the Live API, served from endpoints in the US and the EU.
What the avatar does
The video is generated together with the audio rather than laid over it afterward. Google says the model takes in audio, a camera feed, and a shared screen at the same time, and it can run tool calls in the background without pausing the conversation. The company pitches it for customer service, guided product walkthroughs, and interactive kiosks.
According to Google, Gemini 3.8 Live understands and speaks 97 languages, detects the language on its own, and keeps the lip-sync intact when a speaker switches languages mid-conversation. The Live API overview page in Google's documentation still lists 24 supported languages, so check the model page before you plan around a less common language.
The documentation lists the model as gemini-3.8-live, reached over a stateful WebSocket connection. Audio goes in as 16 kHz PCM and comes back at 24 kHz, and video input is sent as JPEG frames at one frame per second.
- Every customer can pick from Google's library of preset avatars.
- A custom avatar built from a single reference photo requires enterprise allowlisting and a verification process.
- Google says all generated audio and video carry an imperceptible SynthID watermark.
- Gemini 3.8 Live Extended Thinking, the variant with deeper reasoning, remains in private preview.
What it costs
Google's price list bills the Gemini 3.8 Live API per million tokens on regional endpoints. Text input costs $0.75, image and video input $1.00, and audio input $3.00. Text output costs $4.50, audio output $12.00, and avatar video output $1.00.
The conversion rates decide the bill more than the prices do. One second of audio counts as 25 tokens, and one second of avatar video as 6,192 tokens. Video is charged only while the avatar is speaking, not while it listens.
By our arithmetic, one minute of avatar speech is 371,520 video tokens, or about $0.37, plus about $0.018 for the matching audio, for a total of about $0.39. Video has the lowest per-token price in the table, yet it costs roughly 20 times as much as the audio for the same minute, because it produces about 248 times as many tokens.
Long sessions cost more than that per-minute figure. Google bills each turn for every token in the session context window, so tokens from earlier turns are processed and billed again, up to the context limit you configure. This suggests that for a support agent that talks for twenty minutes, the context cap matters as much as the per-token price.
Who can use it
Access runs through Gemini Enterprise and Google Cloud, not through the consumer Gemini app. Google offers the feature with provisioned throughput and enterprise compliance terms, and points customers to its sales team for custom deployments and avatar allowlisting.
Google has not said whether the SynthID watermark survives when a video calling platform re-encodes the stream. If you plan to rely on the watermark to label synthetic video for your users, that is the question to put to Google before launch.
„Video output charges only apply when the avatar is actively speaking.“
Sources
Related

ChatGPT adds virtual try-on for clothes and keeps your reference photo for later
OpenAI added two shopping features to ChatGPT on October 1, 2026. A Try on button on listings for clothes and accessories generates an image of the item on you from a selfie. The selfie is saved as a reference photo for future try-ons until you change or delete it in settings. A second feature, Favorites, saves products to your ChatGPT Library.
OpenAIverified
Google put its Nano Banana image editor into Docs and Slides and named it Google Pics
Google Pics is a standalone image creation and editing tool built on the Nano Banana model, and the same controls are being wired into Workspace apps. The integration begins with Docs and Slides on September 1, 2026, and reaches Drive in the coming weeks. Access arrives with a Google AI Pro or Ultra subscription and with most Workspace business plans, rolling out over the coming weeks. Google publishes no quota, no separate price, and no list of supported languages.
Googleverified
Gemini Omni 1.1 Flash extends a shot to forty seconds and drafts it in 360p first
Google's video model can now continue an existing clip in ten-second steps up to forty seconds total, reading up to ten seconds of prior footage instead of only the final second. A 360p draft mode runs up to 60 percent faster at a third of the cost, and finished work upscales to 1080p or 4K.
Googleverified


