Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

CLBench-V: New Benchmark Reveals Multimodal Context Learning Still Lags

AI By Crimson AI Hugging Face Papers 30 July 2026 · 00:00 10 views
Share: X Telegram

Hugging Face researchers introduce CLBench-V, a benchmark evaluating multimodal context learning across grounding, new information application, and knowledge acquisition. The best model scores only 0.2847, showing significant room for improvement.

CLBench-V: New Benchmark Reveals Multimodal Context Learning Still Lags

Key points

Hugging Face researchers have released CLBench-V, a comprehensive benchmark designed to evaluate how well multimodal AI models can learn from context—a capability that goes beyond relying solely on pre-trained knowledge. The benchmark addresses a critical gap: existing evaluations focus primarily on textual contexts, whereas real-world tasks often involve multimodal inputs such as figures, tables, maps, and web pages.

CLBench-V organizes tasks along three dimensions: context grounding (linking information to specific parts of the input), new information application (using novel data to solve problems), and new knowledge learning (acquiring entirely new concepts from context). The benchmark combines converted public datasets with newly constructed ones spanning science, finance, long-document understanding, spatial reasoning, and web-based visual question answering.

To reduce construction costs, the team developed automated procedures for generating and filtering domain-specific context-learning tasks. The final benchmark comprises 3,443 instances tested across six recent multimodal models. Results show that the best overall score is only 0.2847, indicating that multimodal context learning is far from saturated.

Among the models evaluated, InternVL3.5-30B-A3B performed best on context grounding and new knowledge learning, while Qwen3.5-Plus excelled at new information application. The paper also analyzes judge reliability, context length, image count, and representative failure cases. Code and datasets are available on GitHub.

ModelOverall ScoreContext GroundingNew Info ApplicationNew Knowledge Learning
InternVL3.5-30B-A3BBest overallBest-Best
Qwen3.5-Plus--Best-
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1