Lexicon · The risk of it

Automated AI R&D

Plain English. AI systems doing the work of AI research: writing training and infrastructure code, generating and refining their own scaffolds, proposing and running experiments, evaluating results. This is the concrete mechanism inside recursive self-improvement claims — RSI is the loop and its consequences; this entry is the machinery. What labs actually report today is automation of parts of the research pipeline, with humans still setting direction and judging results.

Why it moves money. The mechanism decides whether capex compounds. If models genuinely accelerate the research that improves models, a lab's compute converts into capability at an increasing rate and leads widen; if the automation is confined to engineering grunt work — the pattern a Princeton-led study found — the gains are real but ordinary productivity, available to every competitor. Lab-reported numbers (Anthropic's institute reports engineers shipping 8x the code per quarter versus 2021–25) are also recruitment and valuation marketing, and deserve the scrutiny that dual purpose earns.

What to watch. Attribution done properly: what fraction of a reported gain is credited to the model, measured how, against what baseline — and whether the harness itself moves inside the training loop, the most direct public test of the mechanism.

From the signals. Ornith-1.5 puts the harness inside the training loop, on its own numbers. Tencent says Hy4 helped train itself, and puts a reported 31.8% on the throughput.

Further reading. METR, "RE-Bench: Evaluating frontier AI R&D capabilities".

← All terms