Epoch AI Audits 15 AI Benchmarks and Flags Nine as Flawed
Epoch AI has launched an independent audit program called Benchmark Reviews, finding that nine out of 15 popular AI evaluations contain significant flaws that could distort model comparisons.

Epoch AI has introduced Benchmark Reviews, a third-party auditing initiative designed to verify whether AI evaluation systems accurately measure model performance. In its inaugural round of 15 audits, the organization verified only four benchmarks, flagged nine as flawed, and concluded that two lacked sufficient public data for a proper assessment. The four verified evaluations are ExploitBench v0.1, SimpleQA Verified, PostTrainBench v1.1, and WeirdML v2.
The nine benchmarks labeled as flawed include prominent developer-focused tools such as SWE-bench Verified, SWE-Bench Pro, Terminal-Bench 4.0.0, DeepSWE v1.1, and the Berkeley Function Calling Leaderboard v4. Other flawed evaluations are Humanity's Last Exam, HealthBench Professional, TextQuests, and Lech Mazur Writing. Epoch AI triggers a flawed designation if at least 20 percent of sampled tasks contain errors that impact accuracy, or if there is a systemic grading issue. For benchmarks with over 50 tasks, auditors review a random sample of 50, expanding to 100 if the error rate lands between 15 and 25 percent. Smaller benchmarks are evaluated in their entirety.
The audits look for critical issues like invalid tasks, false positives, false negatives, setup bias, silent versioning, and under-elicitation, where restrictive limits prevent a model from showcasing its true capabilities. For practitioners, these findings mean that relying solely on high-level leaderboard scores can be highly misleading. Because errors in benchmarks like SWE-bench Verified or the Berkeley Function Calling Leaderboard can distort comparisons of coding agents and tool-use models, developers must look closely at specific test configurations, harness settings, and version histories before making procurement or deployment decisions.
To avoid conflicts of interest, Epoch AI will not audit its own benchmarks. The organization noted that a separate analysis previously revealed errors in 42 percent of its own FrontierMath v1 problems, highlighting the necessity of independent reviews. Moving forward, Epoch AI plans to select future benchmarks for audit based on their industry impact, reach, and diversity, while also offering confidential reviews for private benchmarks.
This is our own summary of reporting by AlphaSignal



