A Sample, Not a Faculty
Somebody finally checked the exam paper. Researchers re-annotated 5,700 questions across all 57 subjects of MMLU, the general-knowledge test that anchored nearly every model launch for four years, and found that around 6.49% of it contains errors: wrong answer keys, ambiguous phrasing, questions with no correct option at all. In the virology subset, 57% of the questions they examined were flawed. Correcting the mistakes changed the model rankings, which means part of the ordering we had been reading as capability was models agreeing with a marker who was wrong. Alex Williams collected that study and several like it for Communications of the ACM last week, and a smaller detail in his piece is worse. BIG-bench, built by hundreds of researchers, shipped with a canary string: a unique token embedded in the dataset so anyone training a model could filter the benchmark out, and anyone auditing one could check whether they had. When OpenAI ran contamination checks for the GPT-4 report, BIG-bench had been swallowed by the crawler anyway. The model can produce the canary on request.
Nobody decided to cheat there. The pipeline did it by default, because a held-out test set that has sat on GitHub for three years is not held out in any sense that matters. Instrument noise is a third failure of the same sort: the Leaderboard Illusion authors submitted two identical checkpoints of one model to Chatbot Arena under different names and the scores landed 17 points apart, about the size of gap that gets written up as a generational leap. None of this is subtle, and all of it is fixable in principle. Rotate the questions, proofread the keys, publish confidence intervals, stop reporting single runs.
The objection that doesn't dissolve under better hygiene was made by Raji, Bender, Paullada, Denton and Hanna at NeurIPS in 2021, and I think it's correct and mostly ignored. Treating any benchmark as a measure of general ability is a category error, not a calibration problem. A benchmark is a specific, finite, contextual set of tasks. You can make it bigger, cleaner and fresher, and it will still be specific, finite and contextual. No amount of engineering converts a sample into a faculty. So a cleaned-up leaderboard buys you a more honest number about a narrower thing, and the narrowing is the whole content of the result.
Which is why the measurement work I find worth reading isn't the work that claims to have built a better exam. It's the work that says out loud what smaller quantity it is actually reporting.
François Chollet's version, running since 2019, is that intelligence is not a stock of solved problems but the efficiency with which you acquire skill at problems you've never seen. ARC-AGI is built on the distinction: it tests fluid intelligence rather than crystallized, and restricts itself to a small set of Core Knowledge priors so a system can't win by having read more than the person it's compared against. A model that arrives holding task-specific knowledge the human lacks is scoring the cleverness of whoever encoded it. ARC-AGI-3, released in March, pushes the idea about as far as it goes. Agents are dropped into turn-based environments with no instructions, no stated goal and no reward signal, and have to work out what the game is before they can play it. Scoring compares the actions an agent burns against a human baseline rather than counting right answers. Humans solve 100% of the environments; frontier systems, as of March, score below 1%. The report states the scope plainly, fluid adaptive efficiency on novel tasks and nothing else, and maintaining it is manual work: each version has been rebuilt to resist the optimisation that ate the last one, so what ARC-AGI offers is a gap its authors keep re-opening by hand.
METR changes the unit rather than the questions. Instead of asking what fraction of a fixed set a model gets right, it asks how long a job can be before the model stops finishing it. The 50% time horizon is the human-expert completion time at which an agent succeeds half the time, fitted across a couple of hundred software tasks with real people timed on the same work. January's update grew the suite from 170 tasks to 228 and moved it to new infrastructure. Measured over the full history the horizon doubles roughly every six and a half months; measured since 2023, about every four; since 2024, under three. The doubling time is itself halving, and that is the finding, not the third significant figure METR attaches to each estimate. I like this number better than any accuracy percentage, partly because hours of human work is a unit a non-specialist can hold, and partly because it fails visibly: a saturating suite runs out of long tasks, a shortage you can see in the task list rather than a ceiling hidden in a percentage.
OpenAI's GDPval asks a third thing, whether the model can produce the actual deliverable. It draws 1,320 tasks from 44 occupations across nine sectors, based on real work products, and has experienced professionals from the matching occupation blind-compare model output against human output. It is explicitly positioned against the exam format, which is a fair criticism arriving from a company that spent years publishing exam scores. Its limit is economic rather than conceptual: expert human grading costs money per item forever, so the property that makes it credible is the property that stops it scaling, and the lab funding it is a lab it evaluates. We've already seen how carefully a scoreboard can be arranged when the same party sets the test and reports the result.
These are not competing answers to one question, and reading them as a leaderboard of leaderboards is the mistake. A system can extend its METR horizon by sustaining longer software tasks while staying useless on GDPval's deliverables, and ARC-AGI has nothing to say about either. Adaptation efficiency, autonomous task duration and occupational output quality are three quantities, not three estimates of one. The argument about what the milestone even is persists partly because people keep expecting one of them to settle it.
What I do when I'm choosing a model for real work is duller than any of this. I keep a small set of tasks drawn from work I actually have: verify nine URLs and report honestly which ones are dead, read a thousand-word draft and find the paragraph that sags, take a photograph and say where the faces are. They never get published, so they can't be trained on, and they measure the only thing I need to know, which is whether this model does my job. Public scores are close to useless for the first and silent about the second. That's a purchasing procedure rather than a theory of intelligence, and I've stopped waiting for anyone to hand me the second one.
Sources:
-
Goodhart's Law Comes for Every Benchmark You Trust — Alex Williams, Communications of the ACM
-
AI and the Everything in the Whole Wide World Benchmark — Raji, Bender, Paullada, Denton and Hanna, NeurIPS Datasets and Benchmarks, 2021
-
What is ARC-AGI? — ARC Prize Foundation
-
ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence — ARC Prize Foundation, March 2026
-
Time Horizon 1.1 — METR
-
Measuring AI Ability to Complete Long Software Tasks — METR, arXiv
-
Measuring the performance of our models on real-world tasks — OpenAI
Filed under AI & machine learning
This post is timestamped using Blockchain technology. Verify