A new benchmark called NARU aims to push the boundaries of long-form video understanding by focusing on narrative evolution and cultural nuance in Japanese media. The benchmark, detailed in a recent paper, addresses a gap in existing evaluations that often overlook the joint assessment of these capabilities, especially in high-context, non-English content.
NARU comprises 1,481 questions grounded in 155 videos totaling 146.8 hours, spanning four narrative and five cultural dimensions. To build this resource, the researchers developed a hierarchical memory-based annotation pipeline that transforms raw video into structured event, narrative, and cultural annotations. Questions are then generated via task-oriented synthesis and iterative shortcut removal, ensuring robustness.
The construction process involved 68 native-speaker annotators across two verification stages, ensuring high-quality ground truth. Evaluations across eight model configurations revealed substantial limitations in both long-range narrative integration and culturally grounded reasoning, highlighting persistent gaps in current multimodal large language models (MLLMs).
By exposing these weaknesses, NARU provides a systematic testing ground for developing MLLMs capable of reliably interpreting long-form, high-context video. The work is supported by Infinimind and is available on Hugging Face.