Research Papers
HumanCLAW: New Benchmark Reveals VLMs Lack Embodied Self-Awareness
A new evaluation framework, HumanCLAW, decouples action decision-making from motor control to measure the 'action intelligence' of...
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM
Hugging Face researchers introduce TurboVLA, a new VLA paradigm that bypasses the LLM-centric pathway, achieving 97.7% success on...
TD-JEPA: Mining Temporal Progress from Offline Logs to Boost Latent World-Model Planning
Researchers propose Temporal-Distance JEPA (TD-JEPA), which mines a directed temporal cost from reward-free demonstration logs to...
Hugging Face Researchers Tackle RL Instability in Small Language Models
A new study from Hugging Face identifies three reproducible failure modes in reinforcement learning for small language models (70M...
OmniDelta: Skill-Driven Budget Allocation Boosts Token Compression in OmniLLMs
Researchers propose OmniDelta, a training-free framework that allocates token budgets across modalities based on query intent and...
Hugging Face Research: RL Boosts Code Optimization Accuracy by Up to 125%
A new Hugging Face paper introduces DMC-Optim and a three-stage RL framework that turns execution time into a learnable signal, im...
Hugging Face Unveils Parallel Decoding Distillation for Fast Image and Video Generation
Researchers at Hugging Face propose Parallel Decoding Distillation (PDD), a trajectory-based distillation method that accelerates...
Modus: A Decoder-Only Model for Any-to-Any Multimodal Generation
Researchers introduce Modus, a decoder-only any-to-any multimodal model that treats all modalities symmetrically, eliminating the...
New Study Reveals Contamination Risks Persist in Dynamic Fact-Checking Benchmarks
A new paper from Hugging Face researchers shows that even dynamic benchmarks designed to avoid contamination still contain 17–29%...
PerceptionBench: New Benchmark Reveals MLLMs Struggle with Atomic Visual Perception
Hugging Face researchers introduce PerceptionBench, a benchmark isolating ten atomic visual perception capabilities. Tests on 16 f...
Hugging Face Unveils Shieldstral: A 3B-Parameter Safety Classifier Outperforming Models 7x Its Size
Hugging Face introduces Shieldstral, a compact 3B-parameter multimodal safety classifier that matches or surpasses models nearly s...
Visual Prompt Engineering Boosts Video Model Reasoning, New Study Finds
A new paper from Hugging Face introduces Visual Prompt Engineering (VIPE), showing that automatically modifying task images can si...