Time-horizon evals
Plain English. Instead of asking how smart a model is, measure the length of task it can complete autonomously — expressed in how long the task takes a skilled human — at a stated reliability. METR's research programme made this the field's working metric, with a reported doubling time of roughly seven months.
Why it moves money. Labour substitution scales with task length, not test scores. A model reliable for five-minute tasks is an autocomplete; reliable for a day-long task, it is headcount. The doubling curve is the closest thing the field has to an AGI speedometer, and agent-company valuations implicitly price its continuation. The counterargument matters equally: reliability compounds against you on long tasks, and some argue the effort to push nines of reliability rises toward an exponential wall.
What to watch. Whether the doubling holds on new task suites built to resist memorisation; the reliability threshold quoted (50 per cent is standard; commercial deployment needs far more); and benchmarks that measure whole jobs rather than excerpts.
From the signals. Two new benchmarks put a number on the autonomy ceiling. The argument that agent autonomy hits an exponential wall, and the objection.
Further reading. METR: Measuring AI ability to complete long tasks.