Lexicon · Running it

Latency vs throughput

Plain English. Latency is how fast one user gets their answer; throughput is how many tokens a fleet produces in total. They trade off: batching many users together raises throughput but makes each individual stream slower. Every "tokens per second" headline is quietly choosing a side.

Why it moves money. Vendors have built whole businesses on each dial — SRAM-heavy chips sell speed for a single stream; hyperscale fleets sell aggregate volume. Headline numbers need unpacking: Nvidia's 3,400 tokens-per-second figure was measured on a small dense model filling a quarter of the rack, a best case for the architecture. And speed changes architecture, not just cost — at 1,000 tokens per second, tight agent loops that were impractical become the default design.

What to watch. For any quoted tokens-per-second figure: which model, what precision, what batch size, single-stream or aggregate. The footnote is the number.

From the signals. The Register unpacked what Nvidia's 3,400 tokens/sec figure actually measures. 1,000 tokens a second on a trillion-parameter model changes the harness, not just the cost. Cerebras listed Qwen 3.8 27B at about 1,500 tokens per second.

← All terms