Lexicon · Capability & training

Sample efficiency

Plain English. How much data a learner needs to gain a unit of capability. Humans are startlingly sample-efficient — a teenager learns to drive in twenty hours; models learn from oceans of examples and still miss things a novice wouldn't. The gap is one of the deepest open problems in the field.

Why it moves money. Sample efficiency decides whether scaling keeps paying. Models have already read most of the useful internet; if capability requires ever more data and the data is exhausted, the scaling thesis hits a wall that compute cannot buy through. Everything currently propping that wall up — synthetic data, RL environments, data-quality curation — is a sample-efficiency workaround, and evidence that data quality now outbuys model tweaks repriced curation from janitorial work to core R&D. It also sets the value of proprietary data: the scarcer useful examples are, the more owning them is a moat.

What to watch. Measured capability gains attributed to data quality versus quantity, and progress in domains where data is inherently scarce — robotics being the canonical case.

From the signals. A small-scale study found data drove 12x of pretraining gains, model tweaks 3.7x. Xiaomi attacked robotics' data bottleneck with embodiment-free pre-training.

← All terms