Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Kimi (Moonshot)

Kimi Releases WorldVQA Benchmark to Test Visual World Knowledge in MLLMs

AI By Crimson AI Kimi Blog 26 July 2026 · 14:32 28 views
Share: X Telegram

Moonshot AI's Kimi introduces WorldVQA, a benchmark evaluating factual visual knowledge in multimodal LLMs, revealing that even top models fall below 50% accuracy on long-tail knowledge.

Kimi Releases WorldVQA Benchmark to Test Visual World Knowledge in MLLMs

Key points

Moonshot AI's Kimi has released WorldVQA, a new benchmark designed to measure the factual correctness of Multimodal Large Language Models (MLLMs) regarding visual world knowledge. The benchmark addresses a critical question: whether models truly recognize specific entities or merely hallucinate based on visual patterns.

WorldVQA comprises 3,500 high-quality image-question pairs, each with a single, verifiable ground-truth answer. The dataset spans nine categories and explicitly separates data into Head (common knowledge) and Tail (rare/long-tail knowledge) to assess performance degradation as knowledge becomes more obscure. All pairs underwent rigorous multi-stage human verification to ensure quality.

Results show that even state-of-the-art models struggle, often achieving below 50% accuracy on long-tail visual knowledge. The benchmark also evaluates calibration—the alignment between model confidence and actual accuracy—using Expected Calibration Error (ECE) and Slope metrics. Kimi-K2.5 achieved the best performance with an ECE of 37.9% and a Slope of 0.550, but all models exhibited overconfidence, concentrating predictions in the 90-100% confidence range.

WorldVQA is open-sourced, with the dataset and evaluation scripts available on GitHub and Hugging Face. The team hopes it will drive progress toward more factually reliable multimodal AI.

MetricIdeal ValueKimi-K2.5
ECE (Expected Calibration Error)0%37.9%
Slope (Weighted Average Slope)1.00.550
Source
Kimi (Moonshot) · Kimi Blog
Related news
Research paper
Kimi (Moonshot) 17 Aug 2026

Moonshot AI Unveils Kimi K2.5: Open-Source Visual Agentic Intelligence with Swarm Capabilities

Moonshot AI introduces Kimi K2.5, a native multimodal open-source model with advanced coding and vision, featuring a self-directed...

22
Research paper
Kimi (Moonshot) 17 Aug 2026

Moonshot AI Unveils WorldVQA Benchmark to Test Visual World Knowledge in Multimodal LLMs

Moonshot AI releases WorldVQA, a benchmark with 3,500 image-question pairs designed to measure factual visual knowledge in multimo...

18
Research paper
Kimi (Moonshot) 17 Aug 2026

Kimi Launches Agent Swarm: 100 AI Agents Self-Organize to Tackle Complex Tasks

Kimi (Moonshot) unveils Agent Swarm, a research preview that lets K2.5 deploy up to 100 parallel sub-agents that self-organize int...

15