Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

RefCaptioner: New Framework Grounds Video Captions to Multiple Reference Images

AI By Crimson AI Hugging Face Papers 1 August 2026 · 00:00 20 views
Share: X Telegram

Hugging Face researchers introduce RefCaptioner, a two-stage post-training framework for multi-reference image-grounded video captioning, along with a new benchmark (MRVBench) and a large training corpus. The model outperforms open-source peers while preserving general captioning ability.

RefCaptioner: New Framework Grounds Video Captions to Multiple Reference Images

Key points

Video captioning has long focused on generating fluent descriptions of what happens in a clip, but existing models struggle to explicitly tie local visual elements to specific reference images. A new research paper from Hugging Face introduces multi-reference image-grounded video captioning, a task that demands factual descriptions with phrase-level grounding to multiple reference images.

To tackle this, the authors propose RefCaptioner, a two-stage post-training framework. It combines mixed-data supervised fine-tuning (SFT) with a novel reinforcement learning objective called Hierarchical Coverage-Discounted GRPO. This approach jointly improves reference selection, phrase-level binding, distractor rejection, and cross-reference consistency, all while preserving the model's general video-captioning capabilities.

To support training, the team constructed a large corpus containing 20,000 videos and 171,354 reference images. They also introduce MRVBench, a benchmark designed to evaluate caption factuality and multi-reference grounding on both real-world and AI-generated videos.

Experiments show that RefCaptioner achieves the best overall performance among open-source models and remains competitive on standard video captioning benchmarks. Human evaluation confirms that annotators prefer its captions, and that they enable more source-faithful video reconstruction with both open-source and proprietary video generators.

Dataset/BenchmarkVideosReference Images
Training Corpus20,000171,354
MRVBenchReal + AI-generatedN/A
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1