Qwen-Audio-3.1
Qwen-Audio-3.1 is Alibaba's five-model audio lineup: upgraded ASR, TTS and Realtime (full-duplex voice) models plus two new ones, ASR-Next for multi-speaker transcription and audio understanding and TTS-Next for generating voice, sound effects and background audio in one pass. It launched with price cuts of up to 95%.
Alibaba QwenSpeechLaunched
Why it isn't ranked
Speech recognition, TTS and realtime voice models: no ranked speech or voice category yet.
Launches outside the ranked categories are tracked, not scored: they never appear on the charts or in the rankings.
Banner
On the “New in past 24 hours” banner for 24 hours after launch, until .
Official page
qwencloud.com ↗Sources
“Meet Qwen-Audio-3.1! ASR, TTS & Realtime are fully upgraded, joined by two new models: TTS-Next for audio creation and ASR-Next for audio understanding.”
“Alibaba's AI team Qwen has released Qwen-Audio-3.1, a lineup of five models for speech recognition (ASR), text-to-speech (TTS), and real-time interaction.”