
Alibaba today released the Qwen-Audio-3.0-ASR-Flash speech recognition model. Simply put, it helps AI better understand specialized terminology.
Qwen-Audio-3.0-ASR-Flash has undergone systematic optimization in three key areas: contextual consistency, industry-specific term recognition, and customizable hotwords. It also features speech polishing capabilities and can directly output structured text.
The research team has continued to identify specialized terminology from fields such as healthcare, IT programming, stocks, and public figures, building a high-quality vocabulary covering multiple industries to improve the model’s recognition of technical terms and uncommon words. In the latest internal evaluation of industry terminology, the new model showed significant improvements in its recognition rate for specialized terms across all industries, with the healthcare scenario surpassing 95.36%.

Currently, the Qwen-Audio-ASR-Flash series has been validated in scenarios such as meeting minutes organization, real-time subtitles, recorded educational content, and intelligent customer service. It previously ranked first globally on the AI evaluation platform Artificial Analysis with an error rate of 1.7%.
Qwen-Audio-3.0-ASR is now available through Alibaba Cloud Bailian, offering three versions:
Qwen-Audio-3.0-ASR-Flash: speech recognition, up to 5 minutes
Qwen-Audio-3.0-ASR-Filetrans: offline file transcription
Qwen-Audio-3.0-ASR-Streaming: real-time speech recognition
The Qwen-Audio-3.0-ASR-Flash trial links are as follows:
Qwen-Audio-3.0-ASR-Flash:
https://help.aliyun.com/zh/model-studio/non-real-time-speech-recognition-for-fun-asr-flash
Qwen-Audio-3.0-ASR-Flash-Filetrans:
https://help.aliyun.com/zh/model-studio/fun-asr-recorded-speech-recognition-http-api
Qwen-Audio-3.0-ASR-Flash-Streaming:
https://help.aliyun.com/zh/model-studio/fun-asr-realtime-websocket-api
