Latest stories
StateM Runtime Boosts GPT-5.6 to 95.3% Accuracy on Terminal-Bench 2.1
A new runtime system called StateM improves long-horizon agent performance without changing model weights, achieving 95.3% raw acc...
GenRouter: Adaptive Routing Cuts Image-Gen Costs by 95%
Hugging Face researchers introduce GenRouter, a unified routing framework that adaptively assigns prompts to optimal agentic image...
Hugging Face Proposes ACID-Compliant Framework for Reliable LLM Agents
A new research paper from Hugging Face introduces 'agentic transactions,' reinterpreting ACID database guarantees for LLM agents,...
UI-Mate: Open-Weight GUI Agent Sets New Benchmarks with In-Context Demonstrations
Hugging Face researchers introduce UI-Mate, a foundation GUI agent that combines environment-grounded training with in-context dem...
Hugging Face Unveils ClawGym II: Black-Box RL Framework for Agent Harness Optimization
A new black-box reinforcement learning framework from Hugging Face enables stable, scalable optimization of general agents through...
HarnessEval-W: Agentic Framework for Transparent World Model Evaluation
Hugging Face researchers introduce HarnessEval-W, an agentified evaluation pipeline that decomposes world-model assessments into v...
VibeWorlding: RL-Trained Open-Source Agents Outperform Closed-Source in 3D World Building
Hugging Face researchers introduce VibeWorlding, a unified framework that benchmarks and trains multimodal agents to construct 3D...
MMDiff: New Framework Lets Researchers Isolate and Control Features in Multimodal AI Models
Researchers introduce MMDiff, a framework using multimodal sparse autoencoders to identify and control specific features in multim...
Hugging Face Study: Optimal Data Repetition Scales Mildly with LLM Size
A new paper from Hugging Face reveals that under proportional scaling of model size and training tokens, the optimal repetition of...
Hugging Face Paper: Claim-Level Verification Boosts Reasoning Efficiency
A new training-free method, Claim-Level Reliability Assessment (CLR), improves LLM reasoning accuracy by verifying critical claims...
Second Thought: Parallel Reasoning During Agent Idle Time Cuts Sequential Decoding by Up to 43%
A new training-free framework from Hugging Face, Second Thought, runs auxiliary reasoning branches in parallel while LLM agents wa...
LLMs Develop Brain-Like Modular Architecture, Study Finds
A new preprint shows that large language models spontaneously develop modular neural architectures mirroring the human brain's spe...