Latest stories
SoftVTBench: New Visuo-Tactile Benchmark Tracks Physical Interaction Quality in Deformable-Object Manipulation
Hugging Face researchers introduce SoftVTBench, a synchronized visuo-tactile dataset and deformation-aware benchmark that evaluate...
Hugging Face Researchers Unveil Attribute-Guided Framework to Scale Creative Writing Data Across 13 Genres
A new framework separates thematic seeds from genre-form controls to generate diverse, high-quality creative writing data, boostin...
AdaPop: Adaptive Popularity Boosts LLM Unlearning, Cutting Leakage by 5x
Hugging Face researchers propose AdaPop, an unlearning method that adapts gradient pressure based on fact popularity, reducing lea...
Looped Language Models Boost Compositional Tool Calling, New Study Finds
A new paper from Hugging Face researchers shows that looped language models improve multi-step, compositional tool use through rec...
FM-Bench: New Benchmark Shows Managerial Behavior, Not Scale, Drives Long-Horizon LLM Agents
A new benchmark from Hugging Face, FM-Bench, tests LLM agents managing a football club over 20 simulated years, revealing that man...
New Diagnostics and Action-Conditioned Objectives Improve Latent World Model Planning
Researchers propose diagnostics to measure whether latent distances reflect true task progress in JEPA-style world models, and int...
Hugging Face Researchers Introduce Top-K Prompting to Boost Diverse Retrosynthesis Predictions
A new study from Hugging Face presents Top-K prompting and plausibility-aware training for single-step retrosynthesis, achieving s...
Stampli Uses OpenAI's Codex and ChatGPT Work to Cut Launch Time from Weeks to Days
Stampli, an accounts payable automation company, leveraged OpenAI's Codex and ChatGPT Work to compress weeks of launch production...
New Data-Free Method Traces Language Model Lineage via Weight Signatures
Researchers introduce a passive, data-free technique that verifies whether open-weight language model checkpoints share ancestry b...
SPADE: Self-Play Framework Lets LLMs Design Their Own Training Environments
Hugging Face researchers introduce SPADE, a self-play reinforcement learning framework where a language model acts as both environ...
SemaPLC: Verification-Gated Agent Harness Boosts PLC Code Generation Reliability
SemaPLC, a new agent harness from Midea AI, validates PLC code through external compilation and live runtime execution, achieving...
Hugging Face Introduces SemComp-Bench to Test Whether Video Generators Actually Complete Tasks
A new benchmark from Hugging Face, SemComp-Bench, evaluates video generation models on semantic task completion, measuring both ou...