Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Ex-Omni-2D: Giving AI Dialogue Models a Visible Presence

AI By Crimson AI Hugging Face Papers 12 August 2026 · 00:00 18 views
Share: X Telegram

Hugging Face researchers introduce Ex-Omni-2D, an omni-modal dialogue framework that generates coordinated text, speech, and expressive avatar video, using a visual thought plan and a distilled streaming video generator for efficiency.

Ex-Omni-2D: Giving AI Dialogue Models a Visible Presence

Key points

Omni-modal dialogue models can understand multimodal inputs and synthesize spoken replies, yet their responses remain visually disembodied. To address this, Hugging Face researchers have introduced Ex-Omni-2D, a framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video.

Given a multimodal query, reference image, and reference audio, the model predicts a structured Visual Thought Plan (VTP) describing scene, emotion, and motion. This plan guides the generation of response text and native multi-codebook speech units, which form a shared acoustic-temporal interface. These units are decoded into speech and aligned online with video frames, enabling the response and avatar pathways to be learned from heterogeneous data sources without requiring large-scale query-text-speech-video supervision.

The framework employs a full-sequence Video Generator as the primary teacher, which is distilled into a few-step block-causal Streaming Student. The student's Prefix Streaming mechanism carries a clean latent across consecutive chunks, reducing cumulative late-chunk degradation. With four-step inference, the complete four-GPU pipeline achieves an end-to-end real-time factor (RTF) of 1.293 at 400×720/720×400 resolution, offering a practical quality-efficiency operating point.

The researchers invite feedback on visual presence, streaming generation, and the future of embodied dialogue systems, highlighting the potential of giving AI models a visible presence in conversations.

MetricValue
Inference steps4
End-to-end RTF1.293
Resolution400×720 / 720×400
GPU count4
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

0
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1