Researchers released Long-Horizon-Terminal-Bench, a benchmark of 46 long-horizon tasks testing AI agents on complex, multi-hour tasks, with the best model achieving only 15.2% success, indicating significant room for improvement.

1 min read

Long-Horizon-Terminal-Bench: New Benchmark Tests AI Agents on Long-Horizon Tasks

Introduction

FAQ

What is Long-Horizon-Terminal-Bench?

It is a new benchmark for testing AI agents on long-horizon tasks lasting minutes to hours, with dense intermediate rewards for partial progress.

How does this benchmark differ from previous ones?

It focuses on long-horizon tasks requiring hundreds of episodes and millions of tokens, providing partial rewards for progress instead of binary final evaluation.

What are the key results of the study?

Best model achieved only 15.2% pass@1 at partial reward threshold of 0.95, with average pass rate of 4.3%, indicating significant room for improvement.

Should MENA teams adopt this benchmark?

Yes, researchers and developers in the region can use it to test and improve their AI agents on complex tasks like data analysis and experiment reproduction.

Source: arXiv cs.AI

AI-assisted content, human-reviewed.