As forecasts of explosive AI progress often hinge on the ability of AI agents to automate AI research, a new study introduces a novel evaluation method called 'shadow evaluations' to measure progress toward AI R&D automation. The method involves giving an agent the central, open-ended research question of a high-quality unpublished paper, and having the paper's original authors grade the agent's output.
The researchers ran shadow evaluations on two unpublished NeurIPS 2026 submissions, providing frontier agents with six days and thousands of dollars in compute. While the agents successfully completed all the engineering without human help, they could not make substantial progress toward answering the research questions. Consequently, both papers were unambiguously rejected by the original authors.
The study identifies five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures.
The authors release the expert reviews, survey responses, agent repositories, and logs to support further research. Their findings provide early evidence that today's agents can handle the engineering aspects of AI research but struggle with critical parts of the research lifecycle, such as formulating creative hypotheses and navigating the iterative process of scientific discovery.