AI startup Tavus says its new Griffin model can already persuade a sizable share of people on live video calls that they are speaking with a real human. In a company-described test, 26 of 54 participants who spoke with Griffin-Lite later said they believed the person on the other end was real, a 48% result.
Tavus unveiled Griffin on October 1 and calls it the first Human Interaction Model, an AI system designed to understand and generate face-to-face conversation. The company says the model pays attention to expressions and pauses, not just spoken words. Its previous system scored 2.4% on the same test.
The version used in the trial was Griffin-Lite. According to Tavus, it interacted with 54 people, and 26 said afterward that they believed their conversation partner was a real person. The older system was tested with 41 people and convinced exactly one of them.
How Tavus ran the test
Tavus says participants were told they would be matched with another person for a one-minute video call about what they were looking forward to this year. Only after the call ended were they asked whether it had crossed their mind that the other side might not be real.
That setup differs from the classic Turing test proposed in 1950 by British mathematician Alan Turing. In the original version, a judge communicates with a hidden human and a hidden machine and tries to determine which is which. In Tavus’ test, nobody on the call was told that a bot might be involved.
The results were published on Tavus’ own research page. Participants were recruited through what the company described as an independent research platform. A community note on X has already pointed out that the findings were not independently verified and did not follow a standard protocol. Tavus says people who became suspicious usually did so within 20 seconds.
How the result compares with other AI work
Efforts to make AI appear human in conversation have been building for some time. A UC San Diego study found that OpenAI’s GPT-4.5 convinced judges it was human in 73% of conversations when it was prompted to act like an introverted, internet-savvy young person. That test, however, was text only.
Griffin adds a face and a voice in real time.
Scores on NVIDIA’s VideoFDB benchmark
Tavus says Griffin-Lite ranks first on NVIDIA’s VideoFDB benchmark, which measures live audio and video conversation.
On the generation track, which scores how natural and expressive a model’s responses are, Griffin-Lite received 3.83. The next-best system scored 2.80, while the human reference came in at 3.92.
On the perception track, which tests whether a model understands what it sees and hears, Griffin-Lite scored 3.73. The strongest baseline scored 3.44, and the human reference reached 4.20. Tavus says NVIDIA ran that evaluation independently.
Real-time interaction and latency
Tavus also says Griffin is full-duplex, meaning it can listen, watch, and speak at the same time, more like a phone call than a walkie-talkie. In a Tavus demo, the model coaches a man through a Rubik’s cube based on what it sees in his hands, then pauses when he goes quiet to think.
On NVIDIA H100 chips, which are commonly used in AI data centers, Tavus says Griffin’s average audio-to-video delay is 0.43 seconds. The company says that is half the latency of the next-fastest method.
Why the claims matter outside AI labs
One reason this matters beyond the tech sector is simple: scammers already operate over video calls.
In January, North Korea-linked hackers used deepfakes on Zoom or Teams calls to pose as trusted contacts. Security researchers attributed that intrusion to BlueNoroff, a subsidiary of Lazarus Group. Victims were persuaded to install malware disguised as an audio fix.
In that report, David Liberman, co-creator of Gonka, a decentralized network for AI computing, said photos and video can no longer be trusted as proof that something is real. The report also noted that those models were not as advanced as Griffin.
Defenses, access limits, and safety work
Companies are already improvising defenses. In 2025, Kraken flagged a suspected North Korean job applicant after its security team asked spontaneous questions, including requests for government ID and the names of local restaurants. The candidate struggled to answer.
Decrypt also wrote that it tried Tavus’ models and found the results disappointing. After additional checking, the outlet said Griffin-Lite is not available to customers and is only being offered to select trusted testers as a research preview. Tavus says it is working on disclosure features and with AI safety organizations.
The company says Griffin still needs safety measures before any public release.
Funding and availability
Tavus raised a $40 million Series B in November 2025, led by CRV. The earlier system that scored 2.4% stitched together three separate models for visuals, dialogue, and perception.
Trusted testers can request access to Griffin-Lite by submitting a form on the Tavus website.

