ByteDance Rolls Out SeedRealtime, a Native Full-Duplex Audio-Video Model Now Live in Doubao

ByteDance Rolls Out SeedRealtime, a Native Full-Duplex Audio-Video Model Now Live in Doubao

N
News Editor
2026-08-05 04:53:00
ByteDance has officially launched SeedRealtime, a native full-duplex audio-video large model, according to a Cailianshe report carried by PANews on Aug 5. The model does not rely on stitching separate speech, vision and text components into a pipeline. Instead, it uses a unified architecture that natively fuses all three modalities. That design lets SeedRealtime process continuous multimodal information streams in real time, enabling back-and-forth interaction that carries audio, visuals and text together. ByteDance describes the resulting experience as one where the model can look, listen and speak at the same time. End-to-end human evaluations show that conversation pacing problems fell by half compared with cascade models. The improvement shows up in how naturally the model times its responses: instances of being cut off mid-sentence, replying only after a long pause, or being triggered by background noise and casual chatter all decreased markedly. In addition, the likelihood of completing a single dialogue turn smoothly and coherently improved by a clear margin. SeedRealtime has now gone fully live on the Doubao app, making the new multimodal interaction capability available to users at scale.

ByteDance has officially launched SeedRealtime, a native full-duplex audio-video large model, according to a Cailianshe report carried by PANews on Aug 5. The model does not rely on stitching separate speech, vision and text components into a pipeline. Instead, it uses a unified architecture that natively fuses all three modalities. That design lets SeedRealtime process continuous multimodal information streams in real time, enabling back-and-forth interaction that carries audio, visuals and text together. ByteDance describes the resulting experience as one where the model can look, listen and speak at the same time.

End-to-end human evaluations show that conversation pacing problems fell by half compared with cascade models. The improvement shows up in how naturally the model times its responses: instances of being cut off mid-sentence, replying only after a long pause, or being triggered by background noise and casual chatter all decreased markedly. In addition, the likelihood of completing a single dialogue turn smoothly and coherently improved by a clear margin.

SeedRealtime has now gone fully live on the Doubao app.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
710

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.