A new arXiv dataset, OpenDiscoveryTrace, captures 558 full AI scientist reasoning trajectories and shows that output-only evaluation hides major behavioral gaps, including Claude Opus 4.6 producing roughly 30x more errors than GPT-5.4.

1 min read

OpenDiscoveryTrace: New Dataset Audits AI Scientist Reasoning, Not Just Outputs

What OpenDiscoveryTrace Introduces

FAQ

What is the OpenDiscoveryTrace dataset?

It is a public dataset of 558 complete AI scientific agent trajectories that records reasoning steps, tool calls, errors, and confidence scores while models execute 124 scientific tasks, rather than evaluating final outputs alone.

How does process evaluation differ from output-only evaluation?

Output-only evaluation measures the final answer, while process evaluation reveals how a model reached it. In this study the three frontier models had similar success rates, yet Claude Opus 4.6 made roughly 30x more errors than GPT-5.4, a gap invisible to output-only scoring.

What is the difference between Claude Opus 4.6 and GPT-5.4 errors?

Most of Claude Opus 4.6's errors (66.7%) were tool misuse, whereas 83.6% of GPT-5.4's errors were reasoning failures, meaning each model has a distinct failure profile requiring different mitigation.

Should MENA AI teams adopt process-level evaluation now?

Yes. Governments and enterprises deploying scientific agents in health, energy, or industry need auditable reliability, and OpenDiscoveryTrace offers a concrete framework for governance and model selection before large-scale rollout.

Source: arXiv cs.AI

AI-assisted content, human-reviewed.