Thinking Machines Lab, led by former OpenAI CTO Mira Murati, has released a research preview of its new “Interactive Models.” This innovative technology introduces full-duplex voice AI that can simultaneously listen and respond, breaking away from the traditional turn-based architecture of current voice assistants.
What Is Full-Duplex AI?
Most existing voice assistants operate in a half-duplex mode—you speak, wait for a response, then speak again. Full-duplex AI allows the system to continuously process incoming audio while generating output, enabling interruptions, clarifications, and more dynamic interactions. The result is a conversation that feels far more natural, akin to a real phone call.
Performance and Speed
The lab's TML-Interaction-Small model demonstrates a response time of approximately 0.40 seconds, nearly eliminating the awkward pauses typical of current voice AI. This low latency is critical for applications requiring real-time feedback, such as live translation, interactive gaming, or emergency response assistants.
Next Steps and Potential Impact
Currently in a research preview phase, the technology is not yet publicly available. Thinking Machines Lab plans a limited preview for selected researchers in the coming months, with a broader release expected later this year. The company aims to embed full-duplex capabilities directly into its core AI model, potentially transforming everything from customer service chatbots to in-car voice interfaces and personal virtual assistants.
While the early results are promising, real-world performance and user experience remain unproven at scale. Nonetheless, the involvement of former OpenAI CTO Mira Murati adds significant credibility and excitement. If successful, this technology could mark a major milestone in making voice interactions seamless and human-like.

