XavierFok
← all posts

How context windows actually work

2026-08-15 · by Xavier Fok

# How context windows actually work

Every model launch now leads with the same number. Context went from four thousand tokens to thirty two thousand, then to two hundred thousand, and some announcements now claim a million or more. The number is always presented as a straight upgrade, the way more storage on a phone would be. I have worked with these models long enough to read it differently. A bigger window does expand what you can attempt. It also raises what each call costs, stretches how long you wait, and changes how evenly the model attends to the material you gave it. None of that appears in the launch post, so I want to lay it out plainly.

Working memory, with nothing outside it

The closest everyday analogy I have for a context window is working memory. Sit down with a twenty page document and you are only ever really thinking about a few paragraphs at once. The rest waits on the desk until you turn back to it.

A language model lives under a stricter version of that limit. Whatever sits inside the window right now is the entire world. There is no desk. A fact that scrolled out of the window has no faded presence in the background. It is simply absent until something puts it back in.

The window also holds everything at once. Your instructions, the conversation so far, any documents you pasted, and the answer the model is currently producing all share one fixed space, and they compete for it. A long document shrinks the room left for everything else.

The unit is tokens, and tokens are money

Context is measured in tokens rather than words or pages, so the headline number only means something once you know the unit. A token is a chunk of a word. A short common word like "the" usually costs a single token, while longer or rarer words split into two or three pieces. Across ordinary English prose this averages out to roughly three or four characters per token, which works out to about seventy five words for every hundred tokens.

Run that arithmetic on a 128k window and you get somewhere near ninety six thousand words, roughly two full novels sitting in front of the model at the same time.

The same unit appears on the invoice. Providers bill by the token on the way in, and usually separately on the way out. The size of the context you carry is therefore a cost decision as much as a technical one. A pipeline making dozens of calls a day with a hundred thousand tokens attached to each call runs up a bill that surprises people. I have worked with people who loaded everything they might conceivably need into every single call, then stared at the monthly statement trying to figure out where it came from. A large context should be a deliberate choice for a specific task, never the default setting.

What a big window genuinely buys

I do not want to talk the number down to nothing, because the growth has been real. A large window lets you hand over an entire long document and ask about any part of it without first guessing which pages matter. It lets a long conversation keep its full history instead of dropping the early parts off the edge. For code, it means a model can hold a large codebase and reason about how distant pieces connect. For writing, it means pasting in everything so far and asking for help staying consistent.

Tasks that once needed careful engineering to even attempt now work by default. That part of the story is true. The launch posts just stop there, and the problems begin one step later.

More context means more waiting

Speed degrades with context length in a way demos hide. On a tight prompt, a capable model answers in a second or two. Fill the same model's window with a hundred pages and you can wait many seconds before the first token appears. In a chat, that is an annoyance. In an automated pipeline making hundreds of calls a day, it compounds into a genuine slowdown. Better hardware and smarter attention mechanisms have chipped away at the problem, but more material to process still takes more time. You rarely notice in a one off demo. You notice at volume.

Lost in the middle

The problem I think matters most is also the one discussed least. Models pay uneven attention to their context. Research backs this up and my own experience agrees: material near the beginning or the end of the window gets used well, while material buried in the middle gets underweighted. The failure mode has a name, lost in the middle, and it is documented rather than folklore.

The practical shape of it looks like this. You paste in a fifty page report and ask a question whose answer sits plainly on page twenty seven, and the model gives you a confident wrong answer. Technically it read everything. It just weighted the middle poorly. So the assumption that more context produces better answers breaks down exactly where you were counting on it. A short context built from the right material routinely beats a sprawling one.

The habits I actually use

Three habits shape how I manage context day to day, and together they make a measurable difference to both quality and spend.

The first is aggressive summarization. Instead of carrying the raw conversation history into every new call, I keep a running summary of the decisions and facts that matter, and I send that instead. It is denser, it is cheaper, and it keeps the important points out of the dead zone in the middle of a long context.

People resist summaries because compression loses information, and it does. Look closely at which information, though. The detail that falls out of a good summary is largely the same detail a long context would have underweighted anyway. The decisions, the constraints, the things you keep referring back to, all of that survives compression fine. Most of the time the trade comes out ahead.

The second habit is retrieval instead of stuffing. When I have a large document collection, I search it first and hand the model only the five or ten passages relevant to the current question. Retrieval augmented generation is the formal name for the pattern, and the intuition underneath it is plain: fetch what matters, leave the rest out. Setting up the retrieval step takes some engineering. Once it runs, answers get more reliable, calls get faster, and the cost per call drops noticeably.

The third habit is pruning. Before each call I look at what the context holds and ask what is actually needed right now. Anything that earned its place two steps ago but has nothing to do with the current question gets trimmed. It costs a few extra seconds and it consistently sharpens the answers.

Reading the giant numbers

A two million token window means you could, in theory, load around one and a half million words and ask the model to reason across all of it in one call. Specific tasks genuinely benefit: hunting patterns through a very large codebase, analyzing a year of documents in one pass, working through a book length body of research. For most everyday uses, though, the quality of an answer comes from giving the model the right information, well organized, focused on the question, and positioned where it will actually be attended to. The giant number is a maximum. It says nothing about the optimum.

There is a second thing the headline hides. The published figure is the best case, the theoretical limit under ideal conditions. Effective context, meaning the span over which the model keeps reasoning reliably without losing the middle, usually runs shorter. Some providers publish long context benchmarks, and those are worth hunting down, because a model can claim a 200k window while its retrieval accuracy degrades meaningfully past 50k tokens. The gap between the spec and the benchmark is the real information. A headline number with no benchmark behind it tells me very little, so I go looking for the benchmark at every launch now.

A ceiling on capability

The way I hold all of this together is to treat window size as a ceiling. A larger ceiling gives you more options and lets you attempt work a smaller window could never fit. Having the option available says nothing about whether using the maximum is wise. For most of what I do, a well managed short context with the right material in it beats a giant context stuffed with everything I own. It costs less, it returns faster, and the model weighs it more evenly. When a task truly needs a large context I use one, and I still manage it carefully.

So when the next launch leads with the context window, you know how to read it. Ask whether the model holds its quality across the whole span, what cost and latency look like at the lengths you actually run, and whether your use case needs the size at all. Those questions cut through the announcement cleanly. If you want more plain breakdowns of how this technology behaves in real use, I publish them regularly on [xavierfok.com](/).

Get new guides and videos first — join the Telegram channel.