Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Hugging Face Unveils SwanTale: A Unified Model for Multi-Speaker Speech and Audio Generation

AI By Crimson AI Hugging Face Papers 4 August 2026 · 00:00 37 views
Share: X Telegram

SwanTale, a new model from Hugging Face, unifies zero-shot and instruction-based multi-speaker speech and audio generation, achieving state-of-the-art expressiveness and supporting complex scenes.

Hugging Face Unveils SwanTale: A Unified Model for Multi-Speaker Speech and Audio Generation

Key points

Hugging Face has released a research paper introducing SwanTale, a unified model for multi-speaker speech and audio generation. The model is designed to handle both zero-shot and instruction-based tasks, catering to creators in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production.

In these scenarios, creators often need to design voices without reference recordings, control speaker styles via natural language, and incorporate acoustic scenes with environments and audio effects. SwanTale addresses these needs by supporting both instruct tasks (using a caption of environment, speaker styles, and fine-grained content) and zero-shot tasks (using reference audio alongside the same content).

The work introduces SwanData-Caption, a data pipeline that cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse multi-level captions. On the model side, SwanTale incorporates SwanVAE for high-quality multi-audio-modality generation, along with reward-conditioned quality control, Engram conditioning, and a Unified Mixture-of-Experts (MoE) architecture for multi-task modeling.

Training employs curriculum learning and GRPO post-training to progressively strengthen capabilities. Experimental results show SwanTale leads on multiple key zero-shot and instruct metrics, achieving the best expressiveness scores in both tasks and supporting complex instruct generation involving multi-speaker speech and audio. Demos are available at the project page.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1