Lexicon · Measuring it

Benchmark

Plain English. A fixed set of tasks with a scoring rule, used to compare models: coding tickets, science questions, terminal sessions. The score in the launch deck is a benchmark result, and the benchmark is an instrument — built by someone, maintained by someone, and subject to drift like any instrument.

Why it moves money. Benchmark scores now decide which labs raise capital, which models enterprises buy, and which press releases get written — a lot of weight for instruments that decay. A formal analysis of 60 language-model benchmarks found nearly half saturated, and models are exhausting fixed test sets faster than the evaluation community can replace them (see benchmark saturation). A score without its instrument's condition attached is half a number.

What to watch. The refresh cadence of benchmarks against the release cadence of models, and who curates the tasks — expert curation, not secrecy, is what the evidence says keeps a benchmark alive.

From the signals. Nearly half of 60 benchmarks studied are saturated, and curation is what saves them. Mark's sharper claim: every assessment, human ones included, has been failing since ChatGPT shipped. Goodhart comes for the leaderboards.

Further reading. Artificial Analysis methodology — a working example of benchmark construction and maintenance in public.

← All terms