Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

New Study Reveals Contamination Risks Persist in Dynamic Fact-Checking Benchmarks

AI By Crimson AI Hugging Face Papers 29 July 2026 · 00:00 10 views
Share: X Telegram

A new paper from Hugging Face researchers shows that even dynamic benchmarks designed to avoid contamination still contain 17–29% potentially contaminated claims, inflating Macro-F1 scores by up to 11.34 points.

New Study Reveals Contamination Risks Persist in Dynamic Fact-Checking Benchmarks

Key points

Researchers from Hugging Face have published a critical analysis of contamination in multimodal automated fact-checking (MAFC) benchmarks, accepted at ACM MM 2026. The paper, titled "Novel Claim or Déjà Vu? Rethinking 'Contamination-Free' Dynamic Evaluation for Multimodal Automated Fact-Checking," challenges the assumption that dynamic benchmarks—which use claims published after an LLM's knowledge cutoff—are contamination-free.

The team constructed a new dynamic benchmark, ClaimReview2025Q4, and compared it with the static AVeriTeC benchmark. They developed a Contamination Detection Pipeline to quantify knowledge contamination risk in LLMs and VLMs. Their experiments yielded 16 findings, with three key results: (1) dynamic evaluation reduces but does not eliminate contamination—17.09% to 29.30% of post-cutoff claims remain potentially contaminated; (2) many new claims can be verified using pre-cutoff public knowledge, either directly or by synthesizing multiple pieces; and (3) contamination can inflate Macro-F1 by up to 11.34 points and distort system rankings.

The study provides practical guidelines for trustworthy MAFC evaluation, emphasizing that high benchmark scores do not necessarily reflect real-world fact-checking ability. The authors call for rethinking evaluation practices to focus on true unseen claims.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1