Alibaba Upgrades Fun-ASR With 100ms First-Token Latency and 82.74% Wenzhou Dialect Accuracy

Alibaba Upgrades Fun-ASR With 100ms First-Token Latency and 82.74% Wenzhou Dialect Accuracy

N
News Editor 01
2026-07-22 10:16:13
Alibaba Tongyi Lab said its upgraded Fun-ASR-Realtime cuts first-token latency to the 100-millisecond range and supports 30 languages and 16 dialects, with Wenzhou dialect accuracy reaching 82.74%.
Alibaba Cloudspeech recognitionAI modelsdialect recognition

Alibaba Tongyi Lab said its streaming speech recognition model Fun-ASR-Realtime now delivers first-token latency in the 100-millisecond range, aiming for text output almost as soon as speech ends. The company said recognition accuracy is now close to that of offline models.

Context awareness is the key change

The main upgrade is not just lower latency. Alibaba said the model has stronger context awareness, allowing it to use prior dialogue and real-time hot words for dynamic correction. In one example, the system can revise a mistaken phrase like “Ye Lu” into “Ye Lu,” meaning “night heron,” after later context becomes available.

That matters in streaming ASR because the model must generate text before a full sentence is known. Fast output alone is not enough. Keeping accuracy stable while new context arrives is the harder part.

100-hour island livestream put the model into harsh conditions

Alibaba also pointed to a 100-hour island livestream as a real-world test. During the broadcast, Fun-ASR-Realtime provided live subtitles through outdoor rainstorms and frequent speaker changes, processing more than 60,000 recognized entries and a total of 1.32 million Chinese characters.

The company framed this as evidence of stable recognition in single-channel audio, heavy environmental noise, and overlapping multi-speaker situations. Those conditions are much closer to production use than controlled benchmark settings.

30 languages and 16 dialects supported

Dialect recognition was another focus of the announcement. Fun-ASR-Realtime currently supports 30 languages and 16 dialects, with an average character accuracy rate of 88.62% across dialect testing. Shanghai dialect reached 92.41%, while Wenzhou dialect came in at 82.74%.

Wenzhou speech is often treated as one of the harder Chinese dialect varieties for recognition systems. Alibaba’s published result suggests a stronger level of generalization on lower-resource dialect data.

Offline model Fun-ASR-Flash also posted a benchmark result

Alibaba said its offline model Fun-ASR-Flash ranked first on the word error rate leaderboard of the global AI evaluation platform Artificial Analysis. The company presented that result alongside the streaming model update.

API access for both Fun-ASR-Realtime and Fun-ASR-Flash is now available on Alibaba Cloud Bailian. The underlying open-source toolkit and models can be accessed through ModelScope and GitHub.

This article was originally published by Bit.Fan. For more cryptocurrency news and market insights, visit www.bit.fan.
100

Disclaimer:

The market information, project data, and third-party content displayed on this platform are for industry information sharing only and do not constitute any form of investment advice or return commitment.

Cryptocurrency trading carries high risks. Users should fully assess their risk tolerance and make independent decisions. All profits, losses, and legal responsibilities are borne by the users themselves.