Every word AI reads gets broken into small pieces - and you pay for each piece. Here's why the machine sometimes 'forgets' what you said at the start.
Picture a taxi meter. It doesn't measure by the full kilometer - it ticks on every small stretch, and you pay for each stretch. A token is that small stretch, just for words. It's not a whole word, and it's not a letter - somewhere in between. 'Conversation', for example, breaks into two or three pieces, it doesn't stay whole. The model doesn't read sentences - it reads a sequence of these pieces, and the meter ticks on every single one.
The context window is something else - how much of your conversation the meter can see right now, while the driver is driving. Not the whole route since the start of the day, just the most recent stretch. Once it hits the edge of the window, the oldest part falls out of range. Not because it forgot on purpose - the room simply ran out, and something had to go.
So if you're having a long conversation with an AI assistant - say you're planning a wedding, and on the fiftieth message you write 'the venue we picked at the start' - it might not know which venue you mean. Not because it's dumb. Because that message fell out of its window long ago. It only sees the recent stretch of the conversation, not the whole day.
That's exactly why prices are quoted 'per million tokens' - each piece costs a fraction of a cent, and they need to pile them up into a heap for the number to mean anything. It's like being given the meter's price in bulk, for a thousand ticks at once, instead of one at a time.
Here's what doesn't get said
People think they only pay for the answer. That's not how it works. Every time the AI-model answers you, it has effectively 're-read' the whole conversation so far - your questions, its previous answers, everything in the window. The meter runs for the re-reading too, not just the new part. That's why long conversations without a break get expensive on their own - without you doing anything more than continuing to ask.
That's why I cut conversations into pieces, instead of dragging one endless chat on for weeks. It's not a whim - it's how you save money and keep your point clear. An empty window with a clear question does more work than a crammed window where the machine digs through a pile of old replies to find your point.