In a new research paper, a team of scientists from Tsinghua, MIT, Harvard, CMU, the Flatiron Institute, Microsoft Research, and other institutions introduces ASI-Bench, a benchmark designed to measure AI systems' capacity for autonomous scientific discovery. Unlike existing benchmarks that focus on applying learned knowledge, ASI-Bench evaluates whether AI can explore the unknown, create new knowledge, and produce verifiable results with minimal human intervention.
The benchmark comprises 60 project-level research tasks across 11 scientific domains, built by over 40 experts with more than 31,000 human hours of effort. A unique feature is its B1 → B4 guidance gradient, which progressively removes methodological guidance within the same research project. This tests whether AI can independently select methods, conduct research, and deliver results as guidance fades.
Evaluating 18 state-of-the-art agent–model configurations, the researchers found a dramatic decline in performance when guidance is removed. The average score drops from 50.91 with full methodological guidance to 29.10 when only the method is specified, and further to 26.62 when agents must determine the method themselves. Even the best system only reached 51.60 under autonomous research settings.
These results highlight that current AI systems remain heavily dependent on human guidance and are far from conducting end-to-end, project-level scientific research autonomously. ASI-Bench is open to the public, and the team invites researchers worldwide to contribute new tasks and help accelerate progress toward artificial superintelligence.