Instruction-based video editing has advanced rapidly, but most existing benchmarks focus on visual changes in silent clips or isolated audio edits. Real-world videos, however, contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. This gap is now addressed by AVE-Compass, a new benchmark introduced by researchers for holistic evaluation of audio-visual editing abilities.
AVE-Compass comprises 145 curated source videos, 196 audio-visually coupled editing instructions, and 2,688 fine-grained checklist items. It evaluates four key dimensions: Instruction Following, Fidelity Preserving, Realism, and Editing Intent. The evaluation uses checklist-based multimodal large language model (MLLM) judging and a dedicated realism rubric, complemented by automated cross-modal, video, and audio metrics.
Extensive evaluation reveals that state-of-the-art models still struggle to execute cross-modal instructions while preserving non-target content. To address this, the researchers propose AVE-Agent, a modular agent framework that decomposes complex instructions into dependent subtasks and iteratively improves editing results through self-reflection and evaluator feedback. AVE-Agent demonstrates improvements in instruction execution, fidelity preserving, and audio-visual alignment in joint editing while maintaining competitive perceptual quality.
The benchmark provides a realistic and comprehensive testbed for future research, pushing the field toward more capable audio-visual editing systems.