Lexicon · The risk of it

Inoculation prompting

Plain English. A training-time technique with a vaccination metaphor: when fine-tuning data would teach an unwanted trait alongside a wanted skill, you add an instruction during training that explicitly requests the bad behaviour. The model attributes the trait to the instruction rather than internalising it as a disposition — and at deployment, without the instruction, the trait largely doesn't appear. Two research groups published the idea independently in late 2025, which is why it surfaced across multiple venues at once.

Why it moves money. Data curation is one of the biggest hidden costs in training, and "tainted data teaches bad habits" is one of the biggest hidden risks — labs have publicly attributed models' darker behavioural priors to what sits in the corpus. A cheap, prompt-level lever that neutralises trait learning without filtering terabytes changes the economics of using messy, real-world and synthetic data. The caveat: results are early, from controlled settings, and nobody has shown it holds at frontier scale under adversarial pressure.

What to watch. Whether frontier labs report using it in production training runs — the technique is cheap enough that silence would be informative — and whether traits suppressed this way stay suppressed under distribution shift.

From the signals. Anthropic publicly attributes evil-acting models to dystopian sci-fi in the training corpus — the class of problem inoculation targets.

Further reading. Wichers et al. (2025); Tan et al. (2025).

← All terms