Lexicon · Measuring it

Shadow evaluations

Plain English. A way to test whether an AI agent can do genuine, open-ended research rather than just checkable engineering tasks. The agent is handed the central research question of a high-quality but unpublished paper and asked to answer it; the paper's own authors then grade the output as they would grade a conference submission. Because the paper is not public, the agent cannot have memorised the answer from training data or found it online.

Why it moves money. The most explosive AI valuations assume recursive self-improvement — AI automating AI research — is close. Shadow evaluations are a direct, hard-to-game test of exactly that claim. In the first published run (a Princeton-led study, arXiv 2607.27191), frontier agents completed all of the engineering unaided but produced papers their authors unambiguously rejected, failing on judgement, creativity and backtracking. If that result holds, the timelines for automated AI R&D — and the capex and equity premia priced on them — are running ahead of the evidence.

What to watch. Whether replications widen the sample beyond two papers and a single model generation, and whether any agent ever clears a shadow evaluation on a paper its authors would actually accept. That crossing, not a rising benchmark score, is the observable that would show open-ended research is being automated.

← All terms