Researchers released Expert-validated STEM QA, a 398-question dataset in physics, chemistry, biology, and math, showing frontier models score below 25%, exposing a critical gap in AI capabilities for scientific research.

1 min read

Expert-validated STEM QA reveals frontier AI models fail physics, chemistry, biology, and math benchmarks

Overview

FAQ

What is the Expert-validated STEM QA dataset?

It is a dataset of 398 questions in physics, chemistry, biology, and mathematics, designed and reviewed by 241 experts to ensure high quality and balanced distribution, aiming to evaluate AI models' capabilities in scientific domains.

How does this compare to previous benchmarks like MMLU or GPQA?

Unlike previous benchmarks that suffer from performance saturation and multiple-choice formats, this dataset focuses on verifiable open-ended questions, providing a more realistic assessment of models' abilities in actual scientific use.

Should MENA research institutions adopt this benchmark?

Yes, it can help evaluate models locally before adoption in scientific research, especially given the low performance that warrants caution.

Source: arXiv cs.AI

AI-assisted content, human-reviewed.