AI startup Tavus on Oct. 1 released a new model called Griffin, describing it as the first "human interaction model" that can understand and generate face-to-face conversation. The company said the system pays attention not only to wording, but also to facial expressions and pauses.
Results from Tavus' video call test
Tavus said that in testing, 26 of 54 participants, or 48%, believed they were speaking to a real person after a one-minute video call with Griffin-Lite. The company's previous-generation system persuaded only 1 of 41 participants, according to Tavus.
Participants were told they would have a one-minute video call with another person and discuss something they were looking forward to this year. Only after the call ended were they asked whether they suspected the other side was not human. Tavus said people who became suspicious usually noticed within 20 seconds.
The result came from Tavus' own research page, with participants recruited through what the company described as an independent research platform. A community note on X said the findings had not been independently verified and did not follow standard protocols.
Benchmark scores and latency
On Nvidia's VideoFDB benchmark, Tavus said Griffin-Lite ranked first. In the generation category, Griffin-Lite scored 3.83, compared with 2.8 for the next-best system and 3.92 for the human reference. In the perception category, it scored 3.73, versus 3.44 for the strongest baseline and 4.2 for the human reference. Tavus said Nvidia conducted the evaluation independently.
Tavus also said Griffin supports full-duplex interaction, allowing it to listen, watch, and speak at the same time. On Nvidia H100 chips, the model posted average audio-video latency of 0.43 seconds.
Funding
Tavus completed a $40 million Series B round in November 2025, led by CRV.

