Pretraining
Plain English. The first and biggest phase of building a model: feeding it a vast corpus of text and having it learn to predict the next token, over and over, until statistical prediction starts to look like knowledge. Everything a model "knows" about the world it mostly learned here.
Why it moves money. Pretraining is where the capital goes — the multi-hundred-million-dollar training runs, the GPU clusters, the power contracts. A lab's pretraining bill is the entry ticket to the frontier, so anything that changes its cost changes who can compete. Efficiency claims are therefore moat claims: if a smaller player can pretrain near-frontier capability cheaply, the incumbents' spend stops being a barrier and starts being a liability.
What to watch. The claimed cost of a frontier-class pretraining run, and whether efficiency recipes survive independent replication. Falling entry costs devalue incumbency; rising ones entrench it.
From the signals. Magic claimed a pretraining recipe roughly 50x more compute-efficient than DeepSeek V4 Pro's — a claim, not yet a replication. A small-scale study found data quality drove 12x of pretraining gains against 3.7x from model tweaks. SemiAnalysis argues sovereign pretraining is becoming a nation-state calculation.