Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

PAST-Bench: New Benchmark Tests Whether Personal AI Agents Really Improve from Experience

AI By Crimson AI Hugging Face Papers 5 August 2026 · 00:00 21 views
Share: X Telegram

Researchers introduce PAST-Bench, a benchmark with 26 scenarios and 204 episodes to isolate whether personal AI agents improve from retained experience, revealing uneven gains and leading to Hermes+, an enhanced agent with targeted interventions.

PAST-Bench: New Benchmark Tests Whether Personal AI Agents Really Improve from Experience

Key points

Recursive self-improvement—the ability of an AI agent to turn accumulated experience into better future behavior—is a cornerstone of advanced agentic systems. Personal AI agents, which retain preferences, task histories, tool routines, and learned skills across sessions, offer a concrete setting to study this capability. Yet, until now, whether retained experience actually improves agents over time has not been systematically tested.

To address this gap, researchers introduce PAST-Bench, a benchmark designed to isolate the effect of retained experience. Each agent runs through ordered sequences of fresh-session tasks under matched conditions that toggle retained experience on and off. The benchmark spans 26 scenarios and 204 episodes across four capability areas: memory, procedural reuse, information gathering, and update.

The study evaluates seven base models and four agent frameworks. Results show that improvement from retained experience is real but uneven across capabilities. Notably, agents with the same headline gain can differ markedly in whether that gain is supported by evidence of the intended save, retrieve, and update pathway.

Guided by these findings, the team developed Hermes+, an extension of the Hermes agent with five targeted interventions across stages of the agent loop. Hermes+ raises the average gain from retained experience and provides clearer pathway evidence, with its strongest improvement on tasks requiring outdated state to be replaced—though the effect remains capability- and model-dependent.

Together, PAST-Bench and Hermes+ provide an evaluation and diagnostic foundation for studying how persistent agents can progress from merely retaining experience to systematically improving through it. The code is available on GitHub.

Benchmark AspectDetails
Scenarios26
Episodes204
CapabilitiesMemory, Procedural Reuse, Information Gathering, Update
Base Models Evaluated7
Agent Frameworks Evaluated4
Interventions in Hermes+5
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1