Gemini 3.8 Live adds avatars, background tools
Google’s Gemini 3.8 Live now combines bidirectional audio/video, synchronized 24 FPS avatars, live visual understanding, multilingual switching, and non-blocking tool calls. It is generally available through Gemini Enterprise, while custom avatars require allowlisting.
Google is pushing real-time agents beyond voice latency into persistent, visually embodied conversations—but enterprise controls and uneven speech quality will decide whether this is production-ready.
- –Background tool execution keeps conversations alive while APIs run, reducing awkward silence and enabling more natural agent workflows.
- –Native 24 FPS lip-synced avatars remove the need for a separate rendering stack, but add visual compute, cost, and UX complexity.
- –Camera and screen understanding unlocks practical use cases such as claims intake, tutoring, support, and guided shopping.
- –Dynamic language switching and custom vocabulary make the model more useful for global and domain-specific deployments.
- –Early practitioner feedback suggests latency is impressive, but robotic TTS remains a serious limitation for phone-based voice agents.
DISCOVERED
1h ago
2026-09-28
PUBLISHED
1h ago
2026-09-28
RELEVANCE
AUTHOR
AI Revolution