Latest stories
SAF-OPD: A Stable Advantage Fusion Framework to Combine RLVR and On-Policy Distillation
Researchers propose Stable Advantage Fusion (SAF) to combine reinforcement learning with verifiable rewards (RLVR) and on-policy d...
N0-TWAM: First Large-Scale Tactile World-Action Model Predicts Touch and Vision for Contact-Rich Robots
Hugging Face researchers introduce N0-TWAM, a tactile-native world-action model that jointly predicts future vision and contact, t...
Weak-to-Strong On-Policy Distillation: Boosting LLMs with Weaker Teachers
A new Hugging Face research paper introduces Weak-to-Strong On-Policy Distillation (W2S-OPD), a method that improves a strong stud...
Mental World Modeling: A New Framework for Predicting Human Decisions
Researchers introduce Mental World Modeling (MWM), a framework that integrates hidden mental states into world models, showing tha...
Hugging Face Study Reveals Scaling Laws for Text Conditioning in Visual Generation
New research from Hugging Face uncovers that diffusion loss scales with structured language in prompts, leading to a system that o...
Meshy T2: Flow Matching Enables Fast Native 3D Mesh Generation
Hugging Face researchers introduce Meshy T2, a flow-matching framework that generates native polygonal meshes with artist-style to...
QQWorld: New Regularization Method Sharpens Latent World Models
Researchers propose QQWorld, a quantile-quantile matching objective that replaces the Epps-Pulley regularizer in latent world mode...
N_0-VTLA: First Tactile-Pretrained VTLA Model Boosts Contact-Rich Robot Manipulation
Hugging Face researchers introduce N_0-VTLA, a vision-tactile-language-action foundation model that integrates tactile sensing at...
New AISPA Framework Audits System Prompts Across 88 Commercial AI Products
Researchers introduce AISPA, a user-centric framework for auditing system prompts in AI applications, revealing wide variation in...
Hugging Face Researchers Unveil RLSVR: A New Paradigm for Self-Improving LLMs on Open-Ended Tasks
A new paper from Hugging Face introduces Reinforcement Learning with Self-Verifiable Rewards (RLSVR), a task-transformation method...
DeepSeek-V4-Flash Goes Official: Agent Benchmarks Surpass V4-Pro-Preview
DeepSeek has launched the official deepseek-v4-flash API in public beta, featuring improved agent capabilities that outperform V4-...
Fairness Pruning: A Surgical Method to Locate Demographic Bias in LLMs
Researchers introduce Fairness Pruning, a lightweight intervention that locates neurons responsible for demographic bias in GLU-ML...