Alibaba on Sept. 19 released Qwen3.8-LiveTranslate, a new real-time simultaneous interpretation model. In a post published by the Qwen team on its official account, the company said the version uses an Interleave architecture and cuts average latency across 60 languages from 2.8 seconds to 2.3 seconds, while also improving fidelity, fluency, and conciseness.
Average latency falls from 2.8 seconds to 2.3 seconds
The central trade-off in simultaneous interpretation is the balance between waiting and accuracy. The longer a model listens, the more complete its understanding of the speaker’s meaning becomes, but that also means users wait longer for the translation.
The metric cited by the company is average latency, or LAAL, which measures how far the translated output lags behind the source on average. In this release, that figure was reduced from 2.8 seconds to 2.3 seconds, a gap of about half a second. In interpretation settings, the company said, that half-second can determine whether the rhythm of a conversation stays intact.
Three new features target multi-speaker conversations
Alibaba highlighted three new capabilities, all aimed at the same type of use case.
- First, real-time speaker separation. The model can identify who is speaking in situations where multiple people are talking, and through more stable voice cloning, it keeps distinct vocal characteristics in each speaker’s translated output.
- Second, synchronized display of the source and translated text, with both languages shown on screen at the same time.
- Third, long-context disambiguation. The model uses earlier parts of the conversation to resolve names and technical terms, helping keep the translation of the same term consistent throughout the session.
These additions are meant to address cases that one-way, sentence-by-sentence translation struggles with: meetings where several people speak, audiences that need to compare the original text with the translation, and repeated proper nouns or specialized terms within a single conversation.
Cloud delivery, with no open weights found
Qwen3.8-LiveTranslate is available through QwenCloud, and the official announcement links to a blog post and the cloud service. ABMedia said it checked Hugging Face and did not find the model weights listed there, which differs from the release approach used for some open-weight versions in the Qwen series.
Text models in the same Qwen 3.8 family can run locally. The report added that Vitalik previously tested one on a laptop, with inference speed at about 33 tokens per second.
This article first appeared on ABMedia.

