RL environments
Plain English. Simulated tasks and workplaces — a fake codebase, a mock browser, a synthetic negotiation — where a model practises and gets graded, so reinforcement learning has somewhere to happen. If verifiable rewards are the engine, environments are the fuel.
Why it moves money. Environments are to this era what training data was to the last one: the scarce input everyone suddenly needs, with a supply chain forming in plain sight. Labs are paying for realistic, hard, gradable task worlds, and a startup category has appeared to sell them — while data that doubles as an environment (recorded gameplay, logged workflows) is being repriced as a moat. The open question is whether environments stay a purchasable commodity or become in-house infrastructure the way frontier data pipelines did.
What to watch. Prices and buyers in the environment supply chain, and whether labs verticalise. Also the quality problem: an agent trained in a shallow environment ships shallow competence, so environment realism becomes a due-diligence question on every "trained with RL" claim.
From the signals. Dan Luu asked why AI labs haven't built RL environments for testing. US$2.3bn went to a lab whose moat is gameplay data.