August 31st.AI Video Inc. Tavus Release of the Real-Time Dialogue Understanding Model Sparrow-2 and its official use in Tavus PAL, API and enterprise platforms. It no longer uses silence to judge whether the user is finished, but at the same time analyses semantics, vocabulary structure, tone, pause, voice and environmental sound, and continues to judge whether the system should listen, wait, speak or continue to speak。

Sparrow-2 handles audio in 10 ms frame flow, with 100 new conversations per second, four times more particle size than the previous generation; 80 ms requires approximately 7 ms and the previous generation approximately 30 ms. Models can distinguish between "um" etc. and real interruptions, wait six to eight seconds while users think, and request a repeat when the sound cannot be reliably identified, rather than continue to answer on the basis of an error。
According to Tavus, the model uses 16,000 sessions of genuine dialogue and 2000 sessions of synthetic dialogue training. In the 575 real-dialogue event tests, the timing failure rate for Sparrow-2 was 2.1%, about a quarter of the best comparison model; in the face of a pause of more than one second, it was 97% that did not talk. The above-mentioned comparison consists of a Tavus self-construction assessment, with LiveKit using only v1-mini, which can run at the end of the device, not its most advanced model。