Tavus Launches Sparrow-2 to Stop Voice AI From Interrupting Pauses

Tavus Launches Sparrow-2 to Stop Voice AI From Interrupting Pauses

N
News Editor
2026-08-30 06:38:05
AI digital human company Tavus has unveiled Sparrow-2, a real-time conversation comprehension model designed to help voice agents decide when to listen, wait, speak, or continue speaking. Unlike traditional systems that filter out background noise, Sparrow-2 jointly processes speech content, tone, pauses, speaker identity, overlapping voices, and ambient noise, updating the conversational state every 10ms. In tests across 575 real conversational nodes, Sparrow-2 achieved a dialogue timing failure rate of just 2.1%, roughly a quarter of the best comparison model. When users paused for over a second, the model waited 97% of the time without assuming the user had finished. Crucially, this patience didn't come from simply delaying responses—both Sparrow-2 and two baseline systems posted a median response time of 680ms when successfully interjecting.

AI digital human company Tavus has released Sparrow-2, a real-time conversation comprehension model that determines whether a voice agent should listen, wait, speak, or continue speaking.

Instead of filtering out background sounds first, the model evaluates speech content, tone, pauses, speaker identity, other people's voices, and ambient noise together, refreshing its understanding of the dialogue state every 10ms. That allows it to handle scenarios where traditional voice AI often stumbles.

Pause for a few seconds to gather your thoughts, and it waits. Say "uh-huh," and it recognizes you're encouraging the AI to keep going—it won't stop cold. If someone else nearby starts talking, it doesn't necessarily treat that as you cutting in. And when the audio is genuinely unclear, it asks you to repeat yourself instead of guessing.

Tavus tested the model across 575 real conversational nodes to see whether the AI would interrupt or miss responses. Sparrow-2 posted a dialogue timing failure rate of just 2.1%, roughly a quarter of the best comparison model's rate.

When users paused for more than one second to think, the model waited 97% of the time rather than assuming the user had finished speaking. That patience isn't bought with a simple response delay. Sparrow-2 and two baseline systems all logged a median response time of 680ms when successfully taking their turn.

Sparrow-2, in short, is better at judging when to hold off and when to actually speak.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
700

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.