Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

EffectLearner: New AI Framework Removes Objects and Their Effects in Video

AI By Crimson AI Hugging Face Papers 8 August 2026 · 00:00 20 views
Share: X Telegram

Hugging Face researchers introduce EffectLearner, a semantic-reasoning framework for video object removal that also eliminates object-induced effects like shadows and reflections, outperforming baselines on ROSE-Bench and new EffectWorld benchmarks.

EffectLearner: New AI Framework Removes Objects and Their Effects in Video

Key points

Video object removal is a challenging task that requires not only deleting the target object but also its induced effects, such as shadows, reflections, and lighting changes, while maintaining temporal coherence. Existing methods often struggle with complex real-world scenes involving compositional effects or weakly correlated effects. To address this, researchers from Hugging Face propose EffectLearner, a framework that combines a vision-language model (VLM)-based Object-Effect Reasoner with a Diffusion Transformer (DiT)-based Video Eraser.

The Reasoner analyzes a target-highlighted video using a structured prompt to extract compact effect-aware context, guiding the Video Eraser to perform comprehensive removal. Motion-aware mask guidance and motion-consistency supervision improve coverage and stability under object motion and evolving scene dynamics. This approach enables the model to handle long-tail physical phenomena and dynamically evolving interactions.

To support training and evaluation, the team introduces EffectWorld, a large-scale paired video dataset specifically designed for effect-aware object removal. It covers compositional effects, weak object-effect correlations, long-tail phenomena, and dynamic challenges. A progressive training curriculum combines common supervision with complex-effect data to fully exploit the framework's capabilities.

On the standard ROSE-Bench benchmark, EffectLearner outperforms existing baselines on most metrics and shows clear advantages on both EffectWorld-Eval and the challenging EffectWorld-Wild, demonstrating its ability to deliver high-quality video object removal in complex real-world scenes.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1