AgentLens: A New Benchmark for Evaluating Interactive Coding Agent Trajectories
What is AgentLens?
FAQ
What is AgentLens?
AgentLens is an open-source benchmark for evaluating interactive coding agents, focusing on analyzing the full trajectory rather than just pass/fail results.
How is AgentLens different from other benchmarks?
Unlike traditional benchmarks that only test whether a task was completed, AgentLens evaluates how the agent follows instructions, uses tools, verifies work, recovers from errors, and interacts with the user.
Can MENA teams use AgentLens?
Yes, the benchmark is open-source on GitHub, allowing any team to run it in their nightly evaluation pipeline to diagnose behavior, compare versions, and catch regressions.
Source: arXiv cs.AI
AI-assisted content, human-reviewed.