Alibaba’s Qwen team has released two new models in the Qwen3-TTS family: Qwen3-TTS-VD-Flash for voice design and Qwen3-TTS-VC-Flash for voice cloning. Both are now accessible via the Qwen API.
Voice Design (Qwen3-TTS-VD-Flash) allows users to define voices through complex natural language instructions, controlling timbre, prosody, emotion, and persona. This enables full control from “what to say” to “how to say it,” freeing users from relying solely on existing voices or presets. On the InstructTTS-Eval benchmark, it significantly outperforms GPT-4o-mini-tts and Mimo-audio-7b-instruct overall, and surpasses Gemini-2.5-pro-preview-tts in role-playing tests.
Voice Cloning (Qwen3-TTS-VC-Flash) supports 3-second voice cloning and can generate speech in 10 languages: Chinese, English, German, Italian, Portuguese, Spanish, Japanese, Korean, French, and Russian. On the MiniMax TTS Multilingual Test Set, its average word error rate (WER) consistently beats MiniMax, ElevenLabs, and GPT-4o-Audio-Preview.
Both models offer highly expressive, humanlike voices that automatically adjust tone and rhythm according to semantic content. They also feature robust text parsing, handling complex and non-standard text formats accurately. Users can create and persistently store custom voice profiles, enabling multi-turn, multi-role dialogues.