Benchmark saturation
Plain English. When leading models all score near the ceiling of a test, the test stops telling them apart. The benchmark is exhausted — which says the measure is used up, not that the field is finished. Labs optimise hard against whatever is being scored, so popular benchmarks saturate fast, sometimes within months of release.
Why it moves money. Capability claims priced into valuations and procurement decisions lean on benchmark scores, and a score on a saturated benchmark mostly measures who tuned against it most recently. The gap between a headline score and delivered performance on real work is where money gets lost — and as public tests exhaust, the ability to evaluate models privately becomes a purchased scarcity of its own.
What to watch. Index maintainers retiring tests (an honest tell that the measure died), divergence between private held-out evaluations and public scores, and score jumps that trace to harness or settings changes rather than new models.
From the signals. Artificial Analysis removed GPQA Diamond from its index as saturated, replacing it with harder, partly private evaluations. A paper finds nearly half of 60 benchmarks saturated — and expert curation, not privacy, is what resists it. As benchmarks saturate, evaluation becomes the product.