Voice and avatar agents are now a setting. Design the disclosure first.
Most product teams file a voice front end under someday: a separate project, with its own vendors and its own risks.
Over nine days this month, Google made it a setting.
What shipped
Google released three pieces on its own blog:
- Gemini 3.8 Live (September 15). A live voice model that detects and switches among 97 supported languages mid-conversation, and runs tools and API calls in the background while it keeps talking.
- Gemini 3.8 text-to-speech (September 23). Voices designed from a written description, and voice replication from a 30-second sample of your own voice or one you have the rights to use. Google says “users must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created.” Voice replication through AI Studio is not offered in Illinois, Texas, the EEA, the UK, Switzerland or India.
- Live Avatar (September 24). Real-time video of a speaking character, available in Gemini Enterprise. Custom avatars made from a reference image are limited to allowlisted enterprise customers for now.
Google says all of this output carries SynthID, its imperceptible watermark for AI-generated audio and video.
So the voice itself is no longer the hard part. What remains is what the person on the other end knows, agrees to and can undo. That is product design.
A watermark is not a disclosure
SynthID helps tools detect AI-generated content after the fact. A customer on a support call cannot hear it, and a watermark does not tell them whether they are talking to a person or whether the voice belongs to someone real.
Google built consent checks into voice replication and kept it out of several regions. Your product will need the same care in its own flows, and nobody will supply it for you. Treat disclosure the way you treat an empty state or an error message: a designed moment with written copy, owned by someone on the product team.
This is general product education, not legal advice. Rules on synthetic voices, recording and consent differ by place, as Google’s list of excluded regions suggests. Check your own situation with counsel.
When talking beats clicking
Before designing the disclosure, decide whether the flow should talk at all. A simple rule:
Use voice or an avatar when the user’s hands or eyes are busy, the task is a back-and-forth with many branches (setup, troubleshooting, a guided walkthrough), or the user is more comfortable in a language your screens do not support.
Keep the screen when the user compares options side by side, enters precise values such as amounts, dates and addresses, reviews something before money or data moves, or is likely to be in a public place.
If a flow is on both lists, talk through it and confirm on screen. Voice gathers the intent, and the screen shows exactly what will happen before it happens.
Checklist before a voice or avatar ships
- The first line says it is AI. Write the exact sentence, and put a visible label on any avatar. Test whether users remember it after the call.
- Every cloned voice has a consent record. Whose voice it is, what they agreed to, where it may be used, and how they withdraw. Store it with the voice, not in someone’s inbox.
- You have a region map. List where the feature is on and where it is off, and decide what users in the excluded places get instead.
- A human is one phrase away. Pick a phrase such as “talk to a person” that always works, in every supported language.
- Background work is spoken aloud. Live models can call tools while they talk. Have the agent say what it is doing (“I’m checking your order”), and require a spoken or on-screen confirmation before anything that moves money or data.
- The user gets a written record. Send a transcript or summary after the call, including any action the agent took.
- Evals include accents and switching. Build a test set from real recordings (with permission) across your top languages, including calls that switch language partway through. Rerun it whenever the model changes.
- Someone owns trust metrics. Track how often users ask for a person, hang up early or complain that they were misled. Watch these as closely as resolution time.
What to do this week
- Pick one flow that could plausibly talk, and run it through the rule above. Write down which parts stay on screen.
- Draft the disclosure sentence and the avatar label, and give both the same review you would give a pricing page.
- Find every place your product already uses a synthetic or recorded voice, and check that a consent record exists for each.
- Record 20 test calls across your top three languages, and keep them as the start of your eval set.
Sources
- Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking · Google · 2026-09-15
- Gemini 3.8 text-to-speech says hello · Google · 2026-09-23
- Introducing Gemini 3.8 Live with Live Avatar · Google · 2026-09-24
Researched and drafted with AI assistance, checked against the sources above.
Run it as a business of one.
Begin →