Tavus Griffin AI Fools 48% of Video Callers Into Thinking It Is Human
A model that generates a full video presence from one still image, at sub-second latency, moves synthetic video past the avatar stage and into something a live call can carry without the other side noticing.
Reporting from 1 source: GIGAZINE.
Tavus announced Griffin, an AI model that handles real-time video and audio conversation. In a research preview test called Griffin-Lite, 26 of 54 participants on a one-minute live video call judged the AI to be a real human, about 48%. Griffin generates 720p video from a single reference image, including arms, hands, chair, shadows, and background, with an average delay of 0.43 seconds on NVIDIA H100 hardware.
Tavus calls Griffin a Human Interaction Model. It reads facial expressions, gaze, tone, and silence at intervals under one second, then generates its own voice, expression, and gestures in parallel rather than waiting for the user to finish speaking. The company splits the system into a Continuous Conversational Modeling engine, which decides what to express and when, and an Audio-Visual Generation engine, which produces the audio and video. Tavus says NVIDIA ran the evaluation independently in September 2026.
Synthesized by Yomimono from the 1 cited source below, including Japanese-language reporting where cited, then editorially reviewed before publishing.