Latest stories
Modus: A Decoder-Only Model for Any-to-Any Multimodal Generation
Researchers introduce Modus, a decoder-only any-to-any multimodal model that treats all modalities symmetrically, eliminating the...
New Study Reveals Contamination Risks Persist in Dynamic Fact-Checking Benchmarks
A new paper from Hugging Face researchers shows that even dynamic benchmarks designed to avoid contamination still contain 17–29%...
PerceptionBench: New Benchmark Reveals MLLMs Struggle with Atomic Visual Perception
Hugging Face researchers introduce PerceptionBench, a benchmark isolating ten atomic visual perception capabilities. Tests on 16 f...
Hugging Face Unveils Shieldstral: A 3B-Parameter Safety Classifier Outperforming Models 7x Its Size
Hugging Face introduces Shieldstral, a compact 3B-parameter multimodal safety classifier that matches or surpasses models nearly s...
Visual Prompt Engineering Boosts Video Model Reasoning, New Study Finds
A new paper from Hugging Face introduces Visual Prompt Engineering (VIPE), showing that automatically modifying task images can si...
Relay-OPD: Fixing Prefix Failure in On-Policy Distillation with Teacher-Student Handoffs
A new method called Relay-OPD detects when a student model goes off track during training and lets the teacher briefly take over,...
CodeNib: Multi-View Data System Boosts Coding Agent Efficiency by Up to 25x
Hugging Face researchers introduce CodeNib, a data system that builds reusable lexical, dense, and structural views per repository...
OpenAI Unveils GPT-5.6: Balancing Frontier Intelligence with Unprecedented Efficiency
OpenAI announces GPT-5.6, a new model that enhances efficiency across inference and agentic workflows, delivering more useful inte...
Hugging Face Unveils Wonder: A Real-Time, Camera-Controllable Video World Model
Wonder is a general-purpose video world model that enables real-time, camera-controllable exploration of generated worlds, support...
Mage-VL: Microsoft's Efficient Streaming Multimodal Model Cuts Visual Tokens by 75%
Mage-VL, a new codec-native streaming foundation model from Microsoft, overcomes Moravec's paradox in vision-language models by us...
Hugging Face Researchers Expose 'Implicit-Association Blind Spot' in AI Agent Memory Systems
A new benchmark, InMind, reveals that state-of-the-art agent memory systems fail to apply stored facts when indirect reasoning is...
ReDesign: Agentic Framework Recovers Editable Design Files from Raster Images
Hugging Face researchers introduce ReDesign, an agentic framework that decomposes raster images into editable layer hierarchies wi...