Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Hugging Face Unveils JoyAI-Video-Edit: Real-Time 720p Video Editing at 30 FPS

AI By Crimson AI Hugging Face Papers 5 August 2026 · 00:00 15 views
Share: X Telegram

JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion model, enables real-time, open-ended video editing without future frames, achieving 720p at ~30 FPS on a single Nvidia B200 GPU.

Hugging Face Unveils JoyAI-Video-Edit: Real-Time 720p Video Editing at 30 FPS

Key points

Hugging Face has released a new research paper introducing JoyAI-Video-Edit, a 16-billion-parameter autoregressive diffusion framework designed for real-time, open-ended video editing. The model operates without access to future frames or a predefined video duration, enabling low-latency causal generation with bounded computational resources.

The framework integrates three key innovations: chunk-wise autoregressive adaptation for streaming, Source-Anchored Distribution Matching Distillation (SA-DMD) to preserve source fidelity during two-step generation, and Long-Horizon Autoregressive Distillation to mitigate accumulated temporal drift. These components collectively reduce train-inference mismatch and ensure long-term temporal consistency.

According to the paper, extensive automatic and human evaluations show that JoyAI-Video-Edit substantially outperforms existing streaming editors and remains competitive with strong offline systems on both short and long videos. The complete system achieves end-to-end 720p video editing at approximately 30 FPS on a single Nvidia B200 GPU.

The code is publicly available on GitHub, allowing researchers and developers to explore and build upon the framework.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1