Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

New Benchmark Reveals Omni-LLMs Struggle as Real-Time Video Assistants

AI By Crimson AI Hugging Face Papers 24 August 2026 · 00:00 14 views
Share: X Telegram

Hugging Face researchers introduce OmniAssistBench, a benchmark for evaluating omni-modal LLMs as interactive video assistants. Results show even top models like Gemini-3-Pro score only 66.4, highlighting major gaps in visual prompts, context retention, and timely responses.

New Benchmark Reveals Omni-LLMs Struggle as Real-Time Video Assistants

Key points

Recent advances in omni-modal large language models (Omni-LLMs) have opened the door to real-time video assistants that can perceive environments and guide users through multi-turn conversations. However, evaluating these assistants is challenging because their unpredictable responses dynamically alter user actions, making static datasets inadequate.

To address this, researchers at Hugging Face introduce OmniAssistBench, a benchmark that reverse-engineers existing Internet videos to simulate continuous interactions. The pipeline deduces logical user goals and segments videos into multi-turn clips, requiring over 1,000 expert person-hours to build.

In tests, proprietary Gemini-3-Pro scored 66.4 out of 100, while open-source Qwen3-Omni-Instruct achieved 51.2. Although models generally understand user inputs, they often provide incorrect or incomplete answers, struggling with visual prompts like hand gestures, failing to maintain historical context, and responding prematurely before target events.

The findings indicate substantial room for improvement before Omni-LLMs can become reliable real-time assistants.

ModelScore (out of 100)
Gemini-3-Pro (proprietary)66.4
Qwen3-Omni-Instruct (open-source)51.2
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 29 Aug 2026

Hugging Face Audit: 110 of 124 AI Evaluations Fail to Support Their Claims

A new commit-bound census of 124 Inspect Evals units reveals that 110 stop before deterministic inference due to missing historica...

4
Research paper
Hugging Face 29 Aug 2026

Aphanta: New Framework Diagnoses When Image Editing Boosts Multimodal Reasoning

Hugging Face researchers introduce Aphanta, a diagnostic framework that evaluates when image-editing intermediates improve multimo...

5
Research paper
Hugging Face 29 Aug 2026

EditaLive! Enables Real-Time Character Video Editing for Live Streaming

Hugging Face researchers introduce EditaLive, a framework for real-time human-centric video editing in live streams, achieving sta...

4