Latest stories
Ouroboros: Self-Developing Coding Agent Sets New Benchmarks in Terminal-Bench and OSWorld
Hugging Face researchers unveil Ouroboros, a self-improving coding agent harness that evolves its own tools and prompts, achieving...
Unsupervised Self-Distillation: LLMs Learn from Their Own Majority Votes
A new Hugging Face paper introduces U-OPSD, a method that lets LLMs self-distill without external supervision by using majority-vo...
Hugging Face's BDH-CQ: 150M Model Sets New Cost-Accuracy Frontier on ARC-AGI-1
A new 150M-parameter reasoning model, BDH-CQ, combines in-context learning with recurrent latent reasoning to achieve 29.5% pass@2...
SPOT: A New Distillation Method to Boost Reasoning in Student Models
Researchers introduce SPOT, a novel on-policy distillation technique that uses sparse probing and outcome calibration to improve r...
Hugging Face Unveils Motif 3: A 314B-Parameter MoE Model with Novel Attention
Hugging Face introduces Motif 3, a 314B-parameter Mixture-of-Experts model with 13.2B active parameters, featuring Grouped Differe...
Sci-VBench: New Benchmark Tests AI Video Generation's Scientific Reasoning
Hugging Face researchers introduce Sci-VBench, a benchmark with 1,253 expert-annotated examples across 60 scientific subjects, rev...
Agent Memory Distillation: Boosting Small LLM Agents with Hierarchical Teacher Memory
A new training-free framework, Agent Memory Distillation (AMD), transfers structured knowledge from large teacher agents to small...
Hugging Face Unveils Macaron-V1: An Open Agent-Model Family for Continual Learning
Hugging Face introduces Macaron-V1, an open agent-model family designed for experiential intelligence, featuring a Mixture-of-LoRA...
Hugging Face Unveils SWE-Bench ProMax: A Harder, Multilingual Benchmark for AI Coding Agents
Hugging Face researchers introduce SWE-Bench ProMax, a multilingual code refactoring benchmark with 170 expert-curated instances a...
OpenAI CFO Shares 5 Lessons for Building an AI-Native Finance Function
OpenAI's CFO Sarah Friar outlines five key lessons for creating an AI-native finance function, covering automated forecasting, str...
OpenAI Pledges Responsible AI Infrastructure in Texas in Letter to Governor Abbott
OpenAI has sent a letter to Texas Governor Greg Abbott, outlining its commitment to responsible AI infrastructure that supports re...
OpenAI's GPT-5.6 Sol Powers Model ML's Finance Workflow
Model ML leverages OpenAI's GPT-5.6 Sol to streamline finance operations, from research and analysis to creating editable PowerPoi...