Lexicon · Running it

Dollars per million tokens

Plain English. The standard unit price for model usage: what an API charges to read (input) or write (output) a million tokens, a token being roughly three-quarters of a word. Output usually costs several times more than input. It is the meter on the inference business.

Why it moves money. The industry's deflation rate lives in this number. Per-token prices for a given capability level have fallen relentlessly, which is wonderful for buyers and brutal for anyone whose revenue model is reselling tokens — volume has to outrun price for revenue to grow. But the headline rate misleads: a model that costs less per token and uses more of them per job can cost more per task, and agents multiply that effect.

What to watch. Price per completed task, not price per token. When a vendor cuts the per-token rate, check whether output lengths, reasoning tokens or retries moved the other way — and whether the cut is permanent or promotional.

From the signals. Silicon Data's token price index fell to a record 97 cents per million in August. OpenAI cut frontier output pricing by a third — and labelled the rate promotional, with an expiry date. GPT-5.5 doubled list pricing while writing 52% longer completions in the most-used context band — roughly a 3x cost multiplier per task there.

← All terms