How do you test whether an AI agent can really do a job?
In 2023 the best model fixed 1.96% of real software bugs. By 2026 the top score on a checked subset was above 80%, yet studies found leaked answers and weak tests behind such scores.
▶ Start the storyYou give it real jobs that people have already done, hide the answer key, and check the result with tests. One example is SWE-bench, which has 2,294 tasks taken from software issues that were actually filed, and actually fixed, in 12 Python repositories. The agent is shown the code as it stood before the fix, plus the text of the issue, and has to write a patch. It never sees the tests that will judge the result.
When it was introduced in October 2023, researchers at Princeton University and the University of Chicago found that the strongest model they tried, Claude 2, fixed only 1.96% of the tasks. By February 2026 the leading published result on a human-checked subset called SWE-bench Verified was 80.9%.
1.96%
Then the benchmark itself was questioned. OpenAI said it would stop quoting the Verified figure: the models being scored had read the repositories the tasks come from, and every frontier model OpenAI tried in a sample could reproduce wording from the problem statements or from the developers' own fix. A separate study, by Reem Aleithan and colleagues, of the system then top of the leaderboard in 2024 found that in 32.67% of its supposedly solved tasks the fix was written out in the issue or its comments, and another 31.08% passed only because the tests were too weak. Take both out and its score dropped from 12.47% to 3.97%.
Wikipedia's article on language model benchmarks calls this Goodhart's law: if models are designed or selected to score highly on a benchmark, it may cease to be a good indicator of model quality. These figures are as of early 2026.
Quiz me
0/3
Recap
A benchmark score can only be trusted as far as the test is unseen and sound.
💡 A trick to remember it · A score is a thermometer: it only tells the truth if the test is not rigged or already seen.
Surprising fact · The first top score in October 2023 was 1.96%; the leading published Verified score reached 80.9% by February 2026, and OpenAI, which had helped build Verified, said it would stop quoting it.
Connects to
- 🔗 How do AI agents plug into tools, and into each other?
- 🔁 How does an AI agent work: think, call a tool, look at the result, repeat?
- 🪤 Why can a web page or an email hijack an AI agent?
- 🎯 How did a 100,000-tests-a-day target make COVID numbers less trustworthy?
- 🦋 Was the first computer bug really a moth?
- Ai agent
Sources (5)
No source, no claim. Every fact in this lesson (15 claims) cites at least one of these.