Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

New NARU Benchmark Tests AI's Grasp of Japanese Long-Form Video Narratives

AI By Crimson AI Hugging Face Papers 22 August 2026 · 00:00 23 views
Share: X Telegram

Researchers introduce NARU, a benchmark with 1,481 questions across 155 Japanese videos (146.8 hours) to evaluate narrative evolution and cultural reasoning in long-form video, revealing significant limitations in current multimodal models.

New NARU Benchmark Tests AI's Grasp of Japanese Long-Form Video Narratives

Key points

A new benchmark called NARU aims to push the boundaries of long-form video understanding by focusing on narrative evolution and cultural nuance in Japanese media. The benchmark, detailed in a recent paper, addresses a gap in existing evaluations that often overlook the joint assessment of these capabilities, especially in high-context, non-English content.

NARU comprises 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To build this resource, the researchers developed a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations. Questions are then generated via task-oriented synthesis and iterative shortcut removal, ensuring robustness.

The construction process involved 68 native-speaker annotators across two verification stages, ensuring high-quality ground truth. Evaluations across eight model configurations revealed substantial limitations in both long-range narrative integration and culturally grounded reasoning, highlighting persistent gaps in current multimodal large language models (MLLMs).

By exposing these weaknesses, NARU provides a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video. The work is supported by Infinimind and is available on Hugging Face.

FeatureValue
Number of questions1,481
Number of videos155
Total video duration146.8 hours
Narrative dimensions4
Cultural dimensions5
Native-speaker annotators68
Model configurations evaluated8
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 29 Aug 2026

Hugging Face Audit: 110 of 124 AI Evaluations Fail to Support Their Claims

A new commit-bound census of 124 Inspect Evals units reveals that 110 stop before deterministic inference due to missing historica...

4
Research paper
Hugging Face 29 Aug 2026

Aphanta: New Framework Diagnoses When Image Editing Boosts Multimodal Reasoning

Hugging Face researchers introduce Aphanta, a diagnostic framework that evaluates when image-editing intermediates improve multimo...

5
Research paper
Hugging Face 29 Aug 2026

EditaLive! Enables Real-Time Character Video Editing for Live Streaming

Hugging Face researchers introduce EditaLive, a framework for real-time human-centric video editing in live streams, achieving sta...

4