Research Papers
MerchantBench: New Benchmark Reveals LLM Agents Struggle with Long-Term E-Commerce Operations
Hugging Face researchers introduce MerchantBench, a 365-day e-commerce simulation that tests LLM agents' long-term coherence. The...
Hunyuan3D-Buffalo 1.0: Unified Multimodal Model for Scalable 3D Generation and Editing
Tencent's Hunyuan3D-Buffalo 1.0 introduces a unified framework for 3D understanding, generation, editing, and part generation, tra...
AURORA-LM: A New Continuous-Latent Diffusion Language Model Outperforms Discrete Token Models
Researchers introduce AURORA-LM, a continuous-latent diffusion language model that preserves high-capacity text latents and learns...
Hugging Face Unveils JoyAI-Video-Edit: Real-Time 720p Video Editing at 30 FPS
JoyAI-Video-Edit, a 16B-parameter autoregressive diffusion model, enables real-time, open-ended video editing without future frame...
LLMs Struggle to Delete Code: New Study Reveals 'Deletion Avoidance' in AI Code Editing
A new study from Hugging Face researchers identifies 'deletion avoidance' in LLM code editing, showing that even top models often...
MemSFT: External Parametric Memory Cuts Alignment Tax in Domain Fine-Tuning
Hugging Face researchers propose MemSFT, a method that uses an external parametric memory to adapt LLMs to specialized domains wit...
Hugging Face Paper: New Framework Transfers Motion Across Morphologically Different Objects
Researchers propose Motion Beyond Morphology (MBM), a two-stage framework that transfers motion between objects with substantially...
GradCuit: New Method Boosts LLM Reasoning by Optimizing Hidden States at Test Time
Researchers introduce GradCuit, a test-time optimization method that directly adjusts latent states in a Transformer layer, achiev...
Hugging Face Unveils WCM: A World Critic Model to Fix Value Estimation in VLA Reinforcement Learning
Researchers at Hugging Face propose the World Critic Model (WCM), a lightweight LeJEPA-based architecture that jointly predicts fu...
SWE-Touch: New Benchmark Reveals Coding Agents Fail When Users Edit Code Mid-Task
Hugging Face researchers introduce SWE-Touch, a benchmark that injects conflicting user edits into coding tasks, showing that even...
Hugging Face Unveils DiffusionGemma: A Diffusion LLM That Generates 1,500 Tokens per Second
DiffusionGemma, an experimental open-weight model from Hugging Face, uses discrete diffusion to generate text in parallel blocks o...
CADENA: A Stepwise AI Approach to Reverse-Engineering CAD Models
Hugging Face researchers introduce CADENA, a model that reconstructs 3D meshes into parametric CAD programs step-by-step, mimickin...