Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Qwen (Alibaba)

Alibaba’s Qwen3-TTS Introduces Voice Cloning and Voice Design Models

AI By Crimson AI Qwen Research 23 December 2025 · 00:00 29 views
Share: X Telegram

Qwen3-TTS family expands with two new models: Qwen3-TTS-VD-Flash for voice design via natural language instructions and Qwen3-TTS-VC-Flash for 3-second voice cloning across 10 languages, both accessible via the Qwen API.

Alibaba’s Qwen3-TTS Introduces Voice Cloning and Voice Design Models

Key points

Alibaba’s Qwen team has released two new models in the Qwen3-TTS family: Qwen3-TTS-VD-Flash for voice design and Qwen3-TTS-VC-Flash for voice cloning. Both are now accessible via the Qwen API.

Voice Design (Qwen3-TTS-VD-Flash) allows users to define voices through complex natural language instructions, controlling timbre, prosody, emotion, and persona. This enables full control from “what to say” to “how to say it,” freeing users from relying solely on existing voices or presets. On the InstructTTS-Eval benchmark, it significantly outperforms GPT-4o-mini-tts and Mimo-audio-7b-instruct overall, and surpasses Gemini-2.5-pro-preview-tts in role-playing tests.

Voice Cloning (Qwen3-TTS-VC-Flash) supports 3-second voice cloning and can generate speech in 10 languages: Chinese, English, German, Italian, Portuguese, Spanish, Japanese, Korean, French, and Russian. On the MiniMax TTS Multilingual Test Set, its average word error rate (WER) consistently beats MiniMax, ElevenLabs, and GPT-4o-Audio-Preview.

Both models offer highly expressive, humanlike voices that automatically adjust tone and rhythm according to semantic content. They also feature robust text parsing, handling complex and non-standard text formats accurately. Users can create and persistently store custom voice profiles, enabling multi-turn, multi-role dialogues.

ModelCapabilityLanguagesBenchmark Performance
Qwen3-TTS-VD-FlashVoice design via natural languageN/AOutperforms GPT-4o-mini-tts, Mimo-audio-7b-instruct on InstructTTS-Eval; surpasses Gemini-2.5-pro-preview-tts on role-playing
Qwen3-TTS-VC-Flash3-second voice cloning10 (Chinese, English, German, Italian, Portuguese, Spanish, Japanese, Korean, French, Russian)Best average WER on MiniMax TTS Multilingual Test Set vs. MiniMax, ElevenLabs, GPT-4o-Audio-Preview
Source
Qwen (Alibaba) · Qwen Research
Related news
Qwen (Alibaba)
Qwen (Alibaba) 3 Aug 2026

Qwen 3.8-Max: Alibaba's 2.4T-Parameter Model Sets New Bar in Autonomous Coding and Real-World Work

Alibaba's Qwen 3.8-Max, a 2.4T-parameter MoE model, demonstrates unprecedented autonomy in coding and work tasks, from building a...

48
Qwen (Alibaba)
Qwen (Alibaba) 19 Mar 2026

Qwen3.5-Max-Preview Debuts on Arena with Strong Preliminary Results

Alibaba's Qwen team has released the preview of Qwen3.5-Max on the Arena platform, showcasing impressive performance in early eval...

27
Research paper
Qwen (Alibaba) 23 Dec 2025

Qwen-Image-Edit-2511: Enhanced Consistency and LoRA Integration

Alibaba's Qwen team releases an improved image editing model with better character consistency, multi-person group photo fusion, b...

30