Lexicon · Capability & training

Self-play

Plain English. A system improving by competing against copies of itself, generating its own opponents and its own training data as it goes. The technique behind AlphaGo's successors, which reached superhuman Go with no human games at all — each generation trains against the last, and the curriculum builds itself.

Why it moves money. Self-play is the proven escape from the human-data ceiling. A learner limited to imitating people tops out near the best person; one that plays itself has no such cap — that is what "superhuman" claims usually rest on mechanically. The constraint is that classic self-play needs a defined game: clear rules, a scoreboard, a symmetric opponent. The live investment question is how far it extends beyond games — into adversarial security training (attacker and defender models sharpening each other), negotiation, and models critiquing their own outputs. Where a domain can be made game-like, the human-data ceiling comes off; where it can't, incumbent data moats survive.

What to watch. Credible self-play results in open-ended domains without crisp win conditions. That is the boundary between a historic games technique and a general capability engine — and it hasn't convincingly moved yet.

Further reading. Silver et al., Mastering the Game of Go Without Human Knowledge (Nature, 2017)

← All terms