Crimson AI NewsA CrimsonLingua Network service
EN ع
← Back to news
Research paper Hugging Face

ASI-Bench: New Benchmark Tests AI's Scientific Autonomy, Reveals Heavy Dependence on Human Guidance

AI By Crimson AI Hugging Face Papers 19 August 2026 · 00:00 8 views
Share: X Telegram

Researchers introduce ASI-Bench, the first benchmark to evaluate AI's innovative exploration and autonomous scientific execution across 60 project-level tasks in 11 domains, showing a sharp performance drop when human guidance is removed.

ASI-Bench: New Benchmark Tests AI's Scientific Autonomy, Reveals Heavy Dependence on Human Guidance

Key points

In a new research paper, a team of scientists from Tsinghua, MIT, Harvard, CMU, the Flatiron Institute, Microsoft Research, and other institutions introduces ASI-Bench, a benchmark designed to measure AI systems' capacity for autonomous scientific discovery. Unlike existing benchmarks that focus on applying learned knowledge, ASI-Bench evaluates whether AI can explore the unknown, create new knowledge, and produce verifiable results with minimal human intervention.

The benchmark comprises 60 project-level research tasks across 11 scientific domains, built by over 40 experts with more than 31,000 human hours of effort. A unique feature is its B1 → B4 guidance gradient, which progressively removes methodological guidance within the same research project. This tests whether AI can independently select methods, conduct research, and deliver results as guidance fades.

Evaluating 18 state-of-the-art agent–model configurations, the researchers found a dramatic decline in performance when guidance is removed. The average score drops from 50.91 with full methodological guidance to 29.10 when only the method is specified, and further to 26.62 when agents must determine the method themselves. Even the best system only reached 51.60 under autonomous research settings.

These results highlight that current AI systems remain heavily dependent on human guidance and are far from conducting end-to-end, project-level scientific research autonomously. ASI-Bench is open to the public, and the team invites researchers worldwide to contribute new tasks and help accelerate progress toward artificial superintelligence.

Guidance LevelAverage Score
Full methodological guidance50.91
Method specified only29.10
No methodological guidance26.62
Best system (autonomous)51.60
Source
Hugging Face · Hugging Face Papers
Related news
Research paper
Hugging Face 29 Aug 2026

Hugging Face Audit: 110 of 124 AI Evaluations Fail to Support Their Claims

A new commit-bound census of 124 Inspect Evals units reveals that 110 stop before deterministic inference due to missing historica...

4
Research paper
Hugging Face 29 Aug 2026

Aphanta: New Framework Diagnoses When Image Editing Boosts Multimodal Reasoning

Hugging Face researchers introduce Aphanta, a diagnostic framework that evaluates when image-editing intermediates improve multimo...

5
Research paper
Hugging Face 29 Aug 2026

EditaLive! Enables Real-Time Character Video Editing for Live Streaming

Hugging Face researchers introduce EditaLive, a framework for real-time human-centric video editing in live streams, achieving sta...

4