Latest stories
Hugging Face Unveils MoE-ViE: Efficient Vision Encoders Outperform Dense Models
Researchers at Hugging Face introduce MoE-ViE, a family of Mixture-of-Experts vision encoders that achieve state-of-the-art perfor...
Hugging Face Researchers Unveil Multi-Byte Prediction to Speed Up Byte-Level Language Models
A new paper from Hugging Face introduces multi-byte prediction (MBP), a method that generates multiple bytes in parallel in byte-l...
LEGO-RL: New Framework Bridges Coding-Agent Harnesses with Scalable Reinforcement Learning
Hugging Face researchers introduce LEGO-RL, a framework that enables policy-gradient training directly on native coding-agent harn...
DiSCO: A Black-Box Defense Against Unsafe Text-to-Image Generation
Researchers propose DiSCO, a zero-shot, black-box defense that optimizes prompts to reduce harmful image generation without modify...
EditBridge: A Diffusion Bridge for Faithful, Efficient 4K Image Editing
Researchers propose EditBridge, a diffusion bridge framework that enables faithful and efficient ultra-high-resolution image editi...
V-RAE: Rethinking Video Latent Spaces for Generation
Hugging Face researchers introduce V-RAE, a video representation autoencoder that builds compact generative latents from frozen vi...
Why Agent Skills Work—and When They Fail: New Study Reveals the Mechanics
A new study from Hugging Face researchers shows that skills enhance LLM agents primarily by stabilizing execution through procedur...
Agent Lightning v1.0: A Lightweight Framework for Harnessed Agentic RL
Hugging Face researchers introduce Agent Lightning v1.0, a compact framework for harnessed agentic RL that improves coding-agent p...
Hugging Face Unveils CoinVE-200K: A New Frontier in Compositional Video Editing
A new dataset, benchmark, and 22B model enable compositional instruction-guided video editing with multi-region attention and temp...
TAMP-Nav: Aligning VLMs with 2D Visual Prompting for Efficient Embodied Navigation
Hugging Face researchers introduce TAMP-Nav, a unified framework that aligns vision-language models with 2D visual prompting, sele...
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements
A new paper from Hugging Face introduces Agentic ESOpt, a framework using evolution strategies for full-parameter fine-tuning of l...
FreeToken: Bringing Frontier-Scale MoE Models to Personal Machines
Hugging Face researchers introduce FreeToken, an edge-native serving system that dynamically maps computation and model state onto...