Scalable oversight and weak-to-strong
Plain English. The research problem of supervising systems more capable than their supervisors — when the model can produce more work, faster, in more domains than any human reviewer can check. "Weak-to-strong generalisation" is the lab framing: can a weaker model's supervision elicit good behaviour from a stronger one, as a stand-in for humans supervising superhuman systems? It is the safety agenda labs most often cite to regulators, and its critics call it unproven at exactly the scale where it matters.
Why it moves money. Oversight is the binding cost on autonomy: every agent deployment carries a supervision budget, and human review does not scale past a demo. The only mechanism that scales is AI overseeing AI — which is already commercial practice, with labs running model-assisted audits over hundreds of millions of transcripts and third parties paid to check the checkers. Whoever makes oversight cheap and credible sells a tax everyone else must pay.
What to watch. Independent-access arrangements — outside evaluators with real transcript access, on terms that survive bad findings — and the recursion problem: oversight exercised by the systems being overseen.
From the signals. Anthropic scanned 481 million transcripts with Claude doing the second-stage review, and METR stepped in with independent access. A position paper argues agent oversight degrades the very human skills it depends on.
Further reading. Burns et al., "Weak-to-Strong Generalization" (OpenAI).