we_are_coded.by CODE · The world, decoded
БГ
Concept

Latency and throughput

The BasicsUpdated on 16 August 2026we are coded

How long you wait versus how many get through overall - two different numbers companies trade off behind your back.

Checked on16 August 2026
In short: latency is how long you wait for the first word of the answer. Throughput is how many answers the system manages to produce per second while serving thousands of people at once. Companies constantly trade one for the other - they sacrifice your waiting time to pack more people together. You only feel one of the two: the seconds before the letters show up.

Latency is the simple thing you feel physically while watching the dots pulse. It's the time from the moment you hit "Send" to the moment the first letter of the answer pops up on screen. Not the whole answer - the first word. Companies measure exactly this number, because it decides whether you stay in the chat or close the tab. That's why most systems "stream" the answer - streaming, sending word by word instead of all at once - to shorten the feeling of waiting, even while the real work behind it keeps going invisibly.

Throughput is a completely different thing, and you never see it directly. It's the number of complete answers the whole system produces per second while thousands of people ask it at once. It's measured in tokens per second - a token, the smallest piece of text the model computes with, part of a word, a character, or a symbol. A GPU, the graphics chip, the actual chip the computation runs on, chews through a fixed amount of work per second, no matter how many people are waiting in the queue in front of it. The more people ask at once, the more that amount gets divided among them.

Here's the trade-off no interface shows you. To raise throughput, engineers group requests into "batches" - batching, bundling several strangers' questions into one shared computation, because the chip chews through a group more cheaply than question by question. The batch saves the company money and electricity. But everyone inside it waits a little longer, because the chip first has to gather enough people before it even starts computing. Your latency falls victim to somebody else's throughput - and nobody asks you.

That's why the same model feels snappy at six in the morning and sluggish at nine at night, without a single line of code changing in between. The model doesn't "think slower" at peak hour. The queue in front of it is longer, the batches heavier, the chips split among more people at once. The speed you feel right now is a snapshot of the current load, not a property of the intelligence on the other end.

The speed of the answer doesn't measure how smart the model is - it measures how many people are waiting behind you in an invisible queue you'll never see.

Sounds good, but

I run a server that hosts my own models, and that's exactly why this difference isn't an abstraction to me. I see it on screen every time I fire up several agents at once instead of one. Every single one's latency spikes, even though the hardware is exactly the same as five minutes ago. It's not a malfunction. The hardware is simply splitting its attention across more tasks at once, and each one waits its turn in the queue.

That's why I don't trust a marketing number like "fastest model". It's measured in lab silence, with one question and an empty queue behind it. The real speed is the one you get on a Thursday night, when everyone is asking at once and the batches are packed to the brim. That's where you can tell who actually built a system for real load, and who built a demo for a screenshot.

The visual is generated code art. No third-party images.
Official primary sources
→MDN Web Docs: Understanding latency (Mozilla)