Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

AVE-Compass: New Benchmark Puts Audio-Visual Video Editing to the Test

AI By Crimson AI Hugging Face Papers 6 August 2026 · 00:00 15 views
Share: X Telegram

Researchers introduce AVE-Compass, a comprehensive benchmark with 145 videos and 2,688 checklist items, revealing that current models struggle with cross-modal editing. They also propose AVE-Agent, a modular framework that improves instruction following and audio-visual alignment.

AVE-Compass: New Benchmark Puts Audio-Visual Video Editing to the Test

Key points

Instruction-based video editing has advanced rapidly, but most existing benchmarks focus on visual changes in silent clips or isolated audio edits. Real-world videos, however, contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. This gap is now addressed by AVE-Compass, a new benchmark introduced by researchers for holistic evaluation of audio-visual editing abilities.

AVE-Compass comprises 145 curated source videos, 196 audio-visually coupled editing instructions, and 2,688 fine-grained checklist items. It evaluates four key dimensions: Instruction Following, Fidelity Preserving, Realism, and Editing Intent. The evaluation uses checklist-based multimodal large language model (MLLM) judging and a dedicated realism rubric, complemented by automated cross-modal, video, and audio metrics.

Extensive evaluation reveals that state-of-the-art models still struggle to execute cross-modal instructions while preserving non-target content. To address this, the researchers propose AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback. AVE-Agent demonstrates improvements in instruction execution, fidelity preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.

The benchmark provides a realistic and comprehensive testbed for future research, pushing the field toward more capable audio-visual editing systems.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1