Skip to content
Model Pareto

Qwen-Audio-3.1

Qwen-Audio-3.1 is Alibaba's five-model audio lineup: upgraded ASR, TTS and Realtime (full-duplex voice) models plus two new ones, ASR-Next for multi-speaker transcription and audio understanding and TTS-Next for generating voice, sound effects and background audio in one pass. It launched with price cuts of up to 95%.
Alibaba QwenSpeechLaunched

Why it isn't ranked

Speech recognition, TTS and realtime voice models: no ranked speech or voice category yet.

Launches outside the ranked categories are tracked, not scored: they never appear on the charts or in the rankings.

Banner

On the “New in past 24 hours” banner for 24 hours after launch, until .

Official page

qwencloud.com ↗

Sources

  • “Meet Qwen-Audio-3.1! ASR, TTS & Realtime are fully upgraded, joined by two new models: TTS-Next for audio creation and ASR-Next for audio understanding.”
    Official@Alibaba_Qwen on XSep 23, 2026x.com ↗
  • “Alibaba's AI team Qwen has released Qwen-Audio-3.1, a lineup of five models for speech recognition (ASR), text-to-speech (TTS), and real-time interaction.”
    RumourThe DecoderSep 23, 2026the-decoder.com ↗

All other launches →