Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

UniSwap: First Streaming Framework for Joint Audio-Visual Identity Swap in Talking Videos

AI By Crimson AI Hugging Face Papers 14 August 2026 · 00:00 12 views
Share: X Telegram

Hugging Face researchers introduce UniSwap, a unified streaming audio-visual diffusion transformer that simultaneously transfers appearance and voice in talking videos, achieving synchronized identity replacement with efficient streaming and stable long-form generation.

UniSwap: First Streaming Framework for Joint Audio-Visual Identity Swap in Talking Videos

Key points

Researchers from Hugging Face have unveiled UniSwap, described as the first framework for streaming joint audio-visual identity replacement in talking videos. The system addresses a key challenge in character replacement: coordinating the transfer of both appearance and voice while preserving the original motion, scene, linguistic content, and audio-video timing.

Unlike existing methods that rely on separately optimized models for visual and audio modalities, UniSwap performs joint transfer within a single audio-visual diffusion transformer. This unified approach ensures multi-modal consistency, which is difficult to achieve with separate models. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre while maintaining the source content and dynamics.

To overcome the scarcity of aligned cross-identity training pairs, the team introduced a swap-and-reconstruct pipeline. This method removes visual and vocal identity from real clips and uses the original clips as reconstruction targets, effectively generating training data without needing paired identities.

The model is built on a bidirectional backbone and progressively adapted through several innovations: In-context Pretraining for joint replacement, Conditional Streaming Adaptation for block-causal KV-cached generation, and an Efficient Self-forcing DMD mechanism that reduces sampling from 30 to just 3 denoising steps per block. Additionally, Efficient Multi-LoRA Switching allows the three DMD roles to share a single frozen backbone, while Feature-RoPE Decomposition keeps cached positions within the training range, enabling stable long-form inference.

Experiments reported in the paper demonstrate strong audio-visual synchronization, competitive identity preservation, efficient streaming, and stable long-form generation, positioning UniSwap as a significant step forward in real-time talking-video editing.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 29 Aug 2026

Hugging Face Audit: 110 of 124 AI Evaluations Fail to Support Their Claims

A new commit-bound census of 124 Inspect Evals units reveals that 110 stop before deterministic inference due to missing historica...

4
Research paper
Hugging Face 29 Aug 2026

Aphanta: New Framework Diagnoses When Image Editing Boosts Multimodal Reasoning

Hugging Face researchers introduce Aphanta, a diagnostic framework that evaluates when image-editing intermediates improve multimo...

5
Research paper
Hugging Face 29 Aug 2026

EditaLive! Enables Real-Time Character Video Editing for Live Streaming

Hugging Face researchers introduce EditaLive, a framework for real-time human-centric video editing in live streams, achieving sta...

4