Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

OmniDelta: Skill-Driven Budget Allocation Boosts Token Compression in OmniLLMs

AI By Crimson AI Hugging Face Papers 29 July 2026 · 00:00 11 views
Share: X Telegram

Researchers propose OmniDelta, a training-free framework that allocates token budgets across modalities based on query intent and content complexity, achieving a new accuracy-efficiency Pareto frontier on audio-video benchmarks.

OmniDelta: Skill-Driven Budget Allocation Boosts Token Compression in OmniLLMs

Key points

Omni-modal Large Language Models (OmniLLMs) that process text, audio, and video simultaneously face significant memory and inference costs due to long token sequences. Existing compression methods typically select important tokens under fixed budgets, but the preceding problem of how to allocate budgets across modalities remains underexplored.

Researchers from Hugging Face introduce OmniDelta, a training-free, skill-driven framework that couples intent-aware inter-modal allocation with content-aware intra-modal allocation. The method constructs audio and video skill pools to shift the fixed retained-token budget according to query demand, then reallocates modality budgets over audio segments and video frames using local complexity and temporal redundancy.

OmniDelta can be combined with existing pruning strategies, preserving the total retained-token ratio while changing where the budget is spent. Experiments on four audio-video benchmarks with two Qwen2.5-Omni models show that OmniDelta establishes a new accuracy-efficiency Pareto frontier across pruning ratios.

At 25% token retention on Qwen2.5-Omni-7B, OmniDelta reduces GPU memory by 22.0% and achieves a 1.64x end-to-end speedup over full-token inference.

MetricValue
Token retention25%
GPU memory reduction22.0%
End-to-end speedup1.64x
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1