Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

MerchantBench: New Benchmark Reveals LLM Agents Struggle with Long-Term E-Commerce Operations

AI By Crimson AI Hugging Face Papers 5 August 2026 · 00:00 19 views
Share: X Telegram

Hugging Face researchers introduce MerchantBench, a 365-day e-commerce simulation that tests LLM agents' long-term coherence. The best LLM achieved only 27.3% of human performance, highlighting a major gap.

MerchantBench: New Benchmark Reveals LLM Agents Struggle with Long-Term E-Commerce Operations

Key points

Large language model (LLM) agents are increasingly deployed as autonomous tool users, but most benchmarks evaluate them on bounded tasks with immediate success criteria. Real-world applications, however, often demand long-term coherence—the ability to maintain purposeful behavior over extended periods while adapting to accumulated evidence. This requires persistent environments where actions constrain future choices, feedback arrives with variable delays, and incoherent behavior leads to measurable cumulative effects.

To address this gap, researchers from Hugging Face and Alibaba introduce MerchantBench, a 365-day, order-level simulation grounded in 98,843 real e-commerce product records. The environment provides agents with 26 tools for product sourcing, listing and pricing control, cash-flow management, and mixed-latency feedback adaptation. Agents must navigate promptly observable upstream supplier events alongside delayed downstream order outcomes, such as refunds, negative reviews, and penalties, forcing them to follow individual order lifecycles and revisit earlier decisions.

The study evaluated eight LLMs under two agent frameworks across 48 year-long runs. The results reveal a substantial gap: the best LLM configuration achieved only 27.3% of the mean final net assets attained by human participants. Analysis shows that long-running does not necessarily mean long-horizon—agents may gradually stop acting, narrow their control loops, fail to respond to delayed feedback, or reinforce incorrect assumptions through memory.

MerchantBench is available with open-source code and a project homepage, providing a valuable resource for future research on long-term agent coherence.

MetricValue
Simulation duration365 days
Real product records98,843
Tools for agents26
LLMs evaluated8
Agent frameworks2
Total runs48
Best LLM performance vs. human27.3% of mean final net assets
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1