How long you wait versus how many get through overall - two different numbers companies trade off behind your back.
Latency is the simple thing you feel physically while watching the dots pulse. It's the time from the moment you hit "Send" to the moment the first letter of the answer pops up on screen. Not the whole answer - the first word. Companies measure exactly this number, because it decides whether you stay in the chat or close the tab. That's why most systems "stream" the answer - streaming, sending word by word instead of all at once - to shorten the feeling of waiting, even while the real work behind it keeps going invisibly.
Throughput is a completely different thing, and you never see it directly. It's the number of complete answers the whole system produces per second while thousands of people ask it at once. It's measured in tokens per second - a token, the smallest piece of text the model computes with, part of a word, a character, or a symbol. A GPU, the graphics chip, the actual chip the computation runs on, chews through a fixed amount of work per second, no matter how many people are waiting in the queue in front of it. The more people ask at once, the more that amount gets divided among them.
Here's the trade-off no interface shows you. To raise throughput, engineers group requests into "batches" - batching, bundling several strangers' questions into one shared computation, because the chip chews through a group more cheaply than question by question. The batch saves the company money and electricity. But everyone inside it waits a little longer, because the chip first has to gather enough people before it even starts computing. Your latency falls victim to somebody else's throughput - and nobody asks you.
That's why the same model feels snappy at six in the morning and sluggish at nine at night, without a single line of code changing in between. The model doesn't "think slower" at peak hour. The queue in front of it is longer, the batches heavier, the chips split among more people at once. The speed you feel right now is a snapshot of the current load, not a property of the intelligence on the other end.
Sounds good, but
I run a server that hosts my own models, and that's exactly why this difference isn't an abstraction to me. I see it on screen every time I fire up several agents at once instead of one. Every single one's latency spikes, even though the hardware is exactly the same as five minutes ago. It's not a malfunction. The hardware is simply splitting its attention across more tasks at once, and each one waits its turn in the queue.
That's why I don't trust a marketing number like "fastest model". It's measured in lab silence, with one question and an empty queue behind it. The real speed is the one you get on a Thursday night, when everyone is asking at once and the batches are packed to the brim. That's where you can tell who actually built a system for real load, and who built a demo for a screenshot.