Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

Shadow Evaluations: A New Test Shows AI Agents Can Engineer but Not Innovate in Research

AI By Crimson AI Hugging Face Papers 30 July 2026 · 00:00 16 views
Share: X Telegram

A new evaluation method, 'shadow evaluations,' reveals that frontier AI agents can complete engineering tasks but fail to make substantial progress on open-ended research questions, as shown by two NeurIPS 2026 case studies.

Shadow Evaluations: A New Test Shows AI Agents Can Engineer but Not Innovate in Research

Key points

As forecasts of explosive AI progress often hinge on the ability of AI agents to automate AI research, a new study introduces a novel evaluation method called 'shadow evaluations' to measure progress toward AI R&D automation. The method involves giving an agent the central, open-ended research question of a high-quality unpublished paper, and having the paper's original authors grade the agent's output.

The researchers ran shadow evaluations on two unpublished NeurIPS 2026 submissions, providing frontier agents with six days and thousands of dollars in compute. While the agents successfully completed all the engineering without human help, they could not make substantial progress toward answering the research questions. Consequently, both papers were unambiguously rejected by the original authors.

The study identifies five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures.

The authors release the expert reviews, survey responses, agent repositories, and logs to support further research. Their findings provide early evidence that today's agents can handle the engineering aspects of AI research but struggle with critical parts of the research lifecycle, such as formulating creative hypotheses and navigating the iterative process of scientific discovery.

Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 31 Aug 2026

Hugging Face Unveils StepGuard: Step-Level Guardrails for Safer AI Agents

StepGuard, a new step-level guard model from Hugging Face, audits agent actions before execution, reducing attack success rates by...

1
Research paper
Hugging Face 31 Aug 2026

Hugging Face Researchers Unveil ABot-Recon for Stable Long-Horizon 3D Reconstruction

ABot-Recon, a new streaming 3D reconstruction model from Hugging Face, achieves stable long-horizon performance using only local t...

1
Research paper
Hugging Face 31 Aug 2026

ContextPilot: Teaching Agents Proactive Context Management via Fine-Grained RL

Hugging Face researchers introduce ContextPilot, a framework that enhances long-horizon agent reasoning by expanding context-editi...

1