Context windows went from 4K tokens to over a million in a few years. Here’s what the number really controls, and where more isn’t always better.
If you’ve ever pasted a long document into an AI chat and been told it’s “too long,” you’ve hit the context window. It’s one of the most talked-about numbers in AI, and also one of the most misunderstood. Here’s what it actually controls.
Think of it as short-term memory, not long-term memory
A context window is everything the model can “see” at once while it’s answering you: your question, any files or previous messages you’ve included, and its own reply so far.
It’s closer to short-term working memory than a saved memory. Once the conversation moves past the window, the earliest parts are quietly dropped — the same way you might forget the start of a long meeting by the time it ends.
Why it’s measured in tokens, not words
Models don’t read in whole words; they break text into smaller chunks called tokens, roughly three-quarters of a word on average in English.
A “128K context window” means about 128,000 tokens, which works out to somewhere around 90,000–100,000 words — a fat novel’s worth of text. Code, non-English languages and unusual formatting tend to use more tokens per word, so the same window holds less of them.
What happens when you go over the limit
If what you send is longer than the window, one of two things happens: the request is rejected outright, or the oldest content gets trimmed automatically to make room.
Either way, the model can no longer “see” what got cut, which is why a long conversation can start forgetting details you mentioned near the beginning.
Bigger isn’t always better
It’s tempting to assume a bigger context window always makes a model smarter. In practice, models still pay more attention to information near the start and end of a huge window than to text buried in the middle — a pattern researchers call the “lost in the middle” effect.
A well-organized 20-page document often gets better answers than a messy 400-page dump, even if the model technically has room for both.