Lexicon · Capability & training

Mid-training

Plain English. Emerging lab jargon for the phase between pretraining and final post-training: continued training on curated, capability-targeted data — long-document corpora to extend context, code and maths to prep for reasoning RL, domain infusions. Less glamorous than either neighbour, increasingly where recipes are won.

Why it moves money. Mid-training is part of the quiet migration of competitive advantage from compute scale to recipe design. Two labs with identical clusters and similar raw corpora can land in very different places depending on what they feed the model in this phase and in what order — which makes recipe secrecy a moat precisely because compute is levelling. It also complicates diligence: "trained on X billion tokens" tells you nothing about the sequencing that made the difference, and phase vocabulary varies enough between labs that stage-by-stage comparisons are mostly guesswork from outside.

What to watch. Open releases that publish full multi-stage recipes, because each disclosure converts a rival's secret into common knowledge — and shortens the half-life of recipe moats generally.

From the signals. Kimi K3 shipped the whole recipe, not just the weights.

← All terms