Stanford Launches Terminal-Bench-Science to Test AI on Real Research Workflows
A new benchmark of 70 expert-curated tasks aims to shift AI evaluation from synthetic tests to authentic scientific practice.
Researchers at Stanford University have launched Terminal-Bench-Science, a continuous benchmark designed to evaluate the ability of AI agents to execute complex scientific research workflows. Developed in collaboration with the Terminal-Bench team and global domain experts, the project seeks to align AI development with the actual needs of practicing scientists.
The initial release consists of 70 tasks curated by experts and drawn from real-world scientific research workflows. Unlike many existing evaluations, these tasks were designed by the scientists themselves rather than by model developers. The project is led by Stanford researchers Steven Dillmann, Sanmi Koyejo, and Ludwig Schmidt, and is hosted by Stanford University and the Laude Institute. Development of the benchmark has been supported by industry sponsors, including Google, Anthropic, and xAI.
Shifting the Standard for AI Capability
Terminal-Bench-Science serves as an extension of the broader Terminal-Bench and Harbor framework efforts. For years, the industry has relied on synthetic or general-purpose benchmarks to claim "scientific capability" in large language models. However, these tests often fail to capture the nuance, iterative nature, and technical complexity of actual laboratory or theoretical research. By shifting the standard to practitioner-defined tasks, the project creates a direct feedback loop between the scientific community and the engineers building the models.
Implications for Scientific Discovery
This shift toward authentic measurement is critical because it provides a more rigorous assessment of how AI agents can assist in genuine discovery. When AI tools are trained and tested on the specific workflows researchers use daily, the resulting agents are more likely to handle the intricacies of real-world data and complex reasoning. This alignment has the potential to accelerate the development of specialized AI tools that can meaningfully reduce the manual burden on scientists, potentially speeding up the pace of innovation across multiple disciplines.
The Path Forward
As a continuous benchmark, Terminal-Bench-Science is intended to evolve alongside both AI capabilities and scientific methodologies. Future iterations are expected to expand the library of curated tasks as more domain experts contribute their workflows. The project now stands as a primary test for whether current AI agents can move beyond simple information retrieval and into the role of active, reliable collaborators in the scientific process.