Distillation
Plain English. Training a cheaper model to imitate an expensive one by learning from its outputs at scale. The student doesn't need the teacher's data or code — just enough of its answers.
Why it moves money. Distillation is why a capability lead may not be a durable moat. Every frontier model exposed through an API is, to some degree, a free teacher for its competitors — which compresses the price premium a leader can charge and shortens the window in which training costs can be recouped. It is the mechanism behind near-frontier models at a tenth of the price, and it is now a geopolitical allegation, not just a technique.
What to watch. Whether labs can actually detect and block it (rate limits, output watermarks, account vetting), and whether the price gap between frontier and near-frontier keeps narrowing anyway. If the gap narrows while usage of the cheap tier grows, distillation is winning regardless of what anyone announces.
From the signals. The NSA, CISA and FBI accused Chinese AI firms of industrial-scale distillation — an accusation, not a court finding. An independent probe moved Qwen 18 points toward GPT-5.5 Pro with a prefill trick, hinting at its teacher. Anthropic closed a documented distillation route for new accounts.