Research Papers
CPI-Bench: New Benchmark Aims to Better Evaluate Real-World Image Editing Models
Hugging Face researchers introduce CPI-Bench, a comprehensive benchmark for real-world image editing that covers multi-image tasks...
Hugging Face Unveils HumanTracker: A Human-Aligned Benchmark for Humanoid Motion Tracking
HumanTracker is a large-scale benchmark with a preference-aligned metric (HumanScore) that evaluates humanoid motion tracking base...
Hugging Face Unveils MobileMem: A Benchmark for Year-Scale On-Device Memory
MobileMem is a new benchmark and framework from Hugging Face for evaluating on-device long-term memory using year-scale, multimoda...
Beyond Final Scores: New Evaluation Framework Reveals AI Agents Are Engineering Optimizers, Not Autonomous Researchers
A systematic evaluation of seven frontier models across 36 long-horizon tasks reveals that AI agents excel at engineering optimiza...
Hugging Face Unveils Mobius-v0: Decoupling Knowledge and Reasoning for Efficient AI
Mobius-v0, a new foundation model architecture from Hugging Face, separates global memory from iterative reasoning modules, achiev...
New Benchmark Reveals AI Video Detectors Fail on Crisis Content
A new benchmark, RA-Bench, shows that current AI-generated video detectors fail to generalize across realistic crisis-related vide...
RibAssist 3D: Selective Biplanar Rib-Fracture Detection and 3D Localization from CT Projections
A new open-source research prototype, RibAssist 3D, pairs fractures across orthogonal CT projections for accurate 3D localization,...
PixSDS: Fixing VAE-Induced Pixel Drift in Latent Score Distillation for Cleaner Text-to-3D
New research from Hugging Face identifies a root cause of noisy artifacts in latent score distillation sampling (SDS) and introduc...
Hybrid Pipeline Cuts Gender Bias in English-Romanian Machine Translation by 40 Points
Researchers propose a hybrid pipeline combining LLM-based gender classification with tag-aware neural translation, improving gende...
TailBooster: New Framework Boosts Extreme-Event Prediction in Air Transport
Hugging Face researchers introduce TailBooster, a dual-layer generative framework that synthesizes operationally valid extreme air...
CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation
A new method, CW-BASS v2, adapts pseudo-label selection to the saturated confidence of foundation-model teachers, improving semi-s...
Inaudible Low-Frequency Attacks Can Cripple Audio-Language Models, New Study Warns
Researchers propose a black-box red-teaming method using inaudible low-frequency waveforms to expose vulnerabilities in audio-lang...