Latest stories
DistilVDR: Compact 524M Visual Document Retriever Achieves Near-Teacher Accuracy via Dual-Student Distillation
Hugging Face researchers introduce DistilVDR, a 524M-parameter visual document retriever distilled from an 8B teacher using cosine...
InSight-doc: Adaptive Visual Perception Cuts Hallucinations and Latency in Long-Document AI
Hugging Face researchers introduce InSight-doc, an agentic framework that dynamically adjusts visual resolution during reasoning,...
Early Pruning Boosts Efficiency in Deep Research Agents, Study Finds
A new study from Hugging Face shows that pruning context early in deep research agents yields the largest efficiency gains, reduci...
UniMoMo: Compressing Recommendation MoE Models with Functional Expert Merging
Hugging Face researchers introduce UniMoMo, a post-training compression framework that converts large recommendation MoE models in...
Hugging Face Researchers Boost Multilingual Translation with Reference-Free RL
A new open-source model family, MiLMMT-46-v1.0, uses reference-free reinforcement learning and checkpoint interpolation to surpass...
Decoding-Level Taboo: A New Stress Test Exposes LLM Fragility Off the Beaten Path
Researchers introduce Decoding-Level Taboo, a runtime logit-space stress test that forces LLMs off their nominal generation paths,...
SkillZip: Compressing Agent Skills Without Evaluation Rollouts
Hugging Face researchers introduce SkillZip, an evaluation-free method that compresses self-evolving agent skills by finding minim...
VibeLifeBench: New Benchmark Shows Frontier AI Agents Fail at Long-Horizon Proactive Tasks
Hugging Face researchers introduce VibeLifeBench, a benchmark of 200 multi-week simulated tasks, revealing that even the best fron...
Latent-to-4D: Direct 4D Generation from Video Diffusion Latents
A new method, Latent-to-4D, enables reusable direct 4D generation from video diffusion latents, bypassing RGB and transferring acr...
Ex-Omni-2D: Giving AI Dialogue Models a Visible Presence
Hugging Face researchers introduce Ex-Omni-2D, an omni-modal dialogue framework that generates coordinated text, speech, and expre...
Hugging Face Unveils Mendel Gödel Machine: Smarter Self-Improving Coding Agents
Researchers introduce Mendel Gödel Machine (MGM), a new framework that accelerates self-improving coding agents by leveraging mult...
AdvFD: New Adversarial Fréchet Loss Boosts Visual Generator Post-Training
Researchers propose Adversarial Fréchet Distance (AdvFD), a novel loss that adds a learnable adversarial feature space to static F...