New-ZZZ
RU / EN
Audio and Speech 23 September 2026

Alibaba Qwen-Audio-3.1: a complete audio stack with emotional response and multi-speaker recognition

N
New-ZZZ desk
X @Alibaba_Qwen · 5 days ago

Alibaba introduced Qwen-Audio-3.1—a comprehensive and fully updated audio stack that covers all stages of sound processing: from recognition to generation. The set includes five models, including the new components TTS-Next and ASR-Next.

The system significantly improved speech recognition (ASR) capabilities by adding multilingual and dialect recognition with automatic polishing, which cleans transcripts of extraneous words and repetitions.

The new models allow for the creation of high-quality content: TTS-Next generates voice, sound effects, and background audio in a single pass, making it ideal for podcasts and audiobooks. Furthermore, the Realtime module simulates natural conversation, responding with empathy and adjusting the speech pace depending on the interlocutor's emotional state. This makes Qwen-Audio-3.1 a powerful tool not only for transcription but also for creating emotionally rich audio content.

Why it matters

  • —Offers a complete audio processing cycle (ASR, TTS, Realtime) in one ecosystem.
  • —Integration of emotional intelligence into real-time interaction, enhancing the quality of user experience.
  • —Improved multi-speaker recognition (ASR-Next) with emotion and sound tags, useful for event localization and QA.

Key facts

  • ASR-Next supports multi-speaker recognition with timestamps and emotion analysis.
  • TTS-Next uses a unified LM and diffusion architecture for generating audiobooks and podcasts.
  • The Realtime module simulates natural dialogue, including emotional response and interruption capability.
  • ASR received a polishing function that automatically removes filler words and repetitions.
  • The entire model line is available with significant discounts.
Read the original →

The full text is in the original source. Here we provide a brief summary and key facts.

/ related