Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

OmniPack: Training-Free Token Compression Boosts Omni-Modal LLM Efficiency

AI By Crimson AI Hugging Face Papers 5 August 2026 · 00:00 15 views
Share: X Telegram

A new training-free framework, OmniPack, coordinates structural and semantic token compression to cut computational costs in omni-modal LLMs while preserving performance, achieving up to 98% performance retention with only 16.7% of original FLOPs.

OmniPack: Training-Free Token Compression Boosts Omni-Modal LLM Efficiency

Key points

Omni-modal large language models (Omni-LLMs) excel at audio-visual understanding but face high computational costs due to long, redundant token sequences. Existing token compression methods often fail at low token budgets, either discarding critical structural evidence or underutilizing query-aware audio-visual collaboration.

To address this, researchers propose OmniPack, a training-free framework that combines structural compression before the LLM with task-relevant semantic refinement inside it. The pre-LLM stage removes redundancy via modality-specific importance, global coverage, and similarity-aware merging, while the inner-LLM stage consolidates diverse representations using textual guidance and audio-visual collaboration.

Experiments across five benchmarks and three Omni-LLM backbones show OmniPack consistently achieves the best performance-efficiency trade-off. Notably, on Qwen2.5-Omni-7B, it retains 98.0% of original performance while cutting FLOPs to 16.7%, and still holds 92.9% performance with only 6.8% of original FLOPs.

The framework is model-agnostic and requires no training, making it a practical solution for efficient deployment of omni-modal models in resource-constrained environments.

MetricValue
Performance retention (FLOPs 16.7%)98.0%
Performance retention (FLOPs 6.8%)92.9%
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1