Model collapse
Plain English. The hypothesis that models trained on the output of earlier models degrade over generations — errors and blandness compounding as synthetic text displaces human writing in the training pool. The core result is real: a 2024 Nature paper showed collapse under recursive training in controlled conditions. Whether it bites in practice is contested, because labs do not train recursively on unfiltered output; they curate, filter and mix synthetic data deliberately.
Why it moves money. If collapse operates at industrial scale, fresh human data appreciates, data-licensing deals become moats, and the open web stops being free feedstock. If it is a managed engineering constraint — the position most lab practice implies — then synthetic data keeps cutting training costs and the moat never forms. Sceptics also lean on collapse as a terminal diagnosis for the industry, which overgeneralises a narrow, real failure mode into a general law.
What to watch. A frontier regression credibly attributed to synthetic training data — none is public to date — and the prices actually paid for verified human corpora, which are the market's live estimate of the risk.
From the signals. Mark's caution that "overgeneralisation of model collapse is an easy mistake to make", in a piece on the loud-skeptic incoherence.
Further reading. Shumailov et al., "AI models collapse when trained on recursively generated data", Nature (2024).