AI digital human company Tavus has released Sparrow-2, a real-time conversation comprehension model that determines whether a voice agent should listen, wait, speak, or continue speaking.
Instead of filtering out background sounds first, the model evaluates speech content, tone, pauses, speaker identity, other people's voices, and ambient noise together, refreshing its understanding of the dialogue state every 10ms. That allows it to handle scenarios where traditional voice AI often stumbles.
Pause for a few seconds to gather your thoughts, and it waits. Say "uh-huh," and it recognizes you're encouraging the AI to keep going—it won't stop cold. If someone else nearby starts talking, it doesn't necessarily treat that as you cutting in. And when the audio is genuinely unclear, it asks you to repeat yourself instead of guessing.
Tavus tested the model across 575 real conversational nodes to see whether the AI would interrupt or miss responses. Sparrow-2 posted a dialogue timing failure rate of just 2.1%, roughly a quarter of the best comparison model's rate.
When users paused for more than one second to think, the model waited 97% of the time rather than assuming the user had finished speaking. That patience isn't bought with a simple response delay. Sparrow-2 and two baseline systems all logged a median response time of 680ms when successfully taking their turn.
Sparrow-2, in short, is better at judging when to hold off and when to actually speak.

