As we saw in “How does an LLM work?”, each request to a language model is initially a self-contained operation. Why, then, does a chat seem as though the machine remembers our conversation? Because the software includes earlier messages with each new request. The context window limits how much the model can process at once. Its size is measured in tokens. [1]

The available context window depends on the model and the server settings. Some models can process around 8,000 tokens at once; others can handle a million or more. The server can further restrict the usable size. The following illustrations show what goes into a context window and how it fills up during a conversation. [1] [2]

1. The first request

At the start, the context contains the system prompt with the basic instructions and your first message. There is no LLM response yet. Even this first request already occupies part of the window. [1]

The context window at the start
Image description as text

The context window starts on the left with the blue system prompt and the gold first request. There is no model response yet. The rest of the bar shows available context. Time runs from left to right. Widths are schematic, not actual token measurements.

2. The conversation grows

The response adds more text. With the next request, the software includes the conversation so far alongside your new message. If the LLM uses tools, their calls and outputs also become part of the context. It grows with each round of conversation. Sending the conversation again does not give the model a lasting memory. [1]

The context grows with each round
Image description as text

The blue system prompt is followed by a gold request and a green response. A second request leads to a green model tool call, then a larger orange tool output and finally a green response. The tool call is model output, not tool output. Available context remains on the right. All widths are schematic, not token measurements.

3. What is still relevant?

Discarded approaches, outdated drafts and questions that have long since been resolved remain in the conversation. They take up space alongside the information needed for the current task. In a long or contradictory conversation, the LLM may overlook relevant details or follow existing instructions less reliably. This does not necessarily happen in every long chat, and the token count is not the only factor. [3]

Not every interim result stays current
Image description as text

The context window still contains the blue system prompt, two conversation rounds with a tool call and a larger tool output, plus another request. The first green response has a white dashed inner outline marking an out-of-date interim result. Newer information can supersede earlier interim results. The old content remains present. The mark indicates neither actual deletion nor automatic general drift. Widths are schematic, not token measurements.

4. The window fills up

The context cannot occupy all the available space if there is still a response to generate. Its tokens need room too. How much space is allowed for that response, and its maximum length, depend on the model and the software. The chat software or harness may also impose its own limit before the model's limit is reached. What happens to the conversation at that point depends on the software. [1] [2]

The next response needs space too
Image description as text

On the left are the blue system prompt and the existing conversation history, including the latest request. This existing content is the input. On the right, a subtly tinted block with a dashed white border illustrates space for the next response, which has not yet been generated. It is not existing model output or an actual size prediction. The schematic shows that the response also has to fit inside the limited context window. No fixed token values are given.

5. The Dumb Zone

This is my name for the point at which the context has become so tangled for my task that the responses become less precise or unusable. Where it starts depends on the task, the content and the server. In my experience, I reach this zone by around 70 to 80 per cent of the available context window at the latest. With larger context windows, I often notice it starting at 50 per cent. These are personal observations, not fixed technical limits. I want to reduce the context before I reach that point. Compaction or a session handoff are two ways to do that.

The Dumb Zone as a personal guide
Image description as text

The blue system prompt is followed by two gold requests, green responses, a green tool call and an orange tool output. Available context remains on the right. Red hatching overlays its right-hand portion to mark the Dumb Zone. Its schematic starting point is a personal guide, not a universal threshold or technical limit. Actual usability depends on factors including the task and mix of context. The hatching is not reserved memory and does not imply automatic deterioration. Widths are schematic, not actual token measurements.

6. Compaction

During compaction, your harness or chat software asks the LLM to summarise the session so far. The software then replaces the old conversation in the context with that summary. This frees up space, and the conversation continues with the condensed context. Some harnesses let you guide compaction with your own prompts. How much of the work so far survives depends on the quality of the summary. [3] [4]

The illustration shows the basic idea. Depending on the harness, recent messages may also be retained alongside the summary. [4]

Compaction within the same session
Image description as text

One context bar shows the same session after compaction. The blue system prompt is retained. A compact purple summary replaces the old conversation history, leaving substantially more available context. No new session is started. Red Dumb Zone hatching remains visible in the right-hand portion of the available context. Its schematic starting point is a personal guide, not a technical limit or memory reservation. All widths are schematic, not token measurements.

7. Session handoff

During a session handoff, the LLM prepares a handover for a new session. It records the current goal, the state of the work, important decisions and the next steps. The new session starts with this handover rather than the previous conversation. The emphasis is less on summarising the conversation and more on making it possible to continue the work. Here too, the result depends on which information is carried over and how reliably the handover describes it. [5]

One practical example is my project pi-blitz-handoff, which organises handoffs between Pi sessions. I introduce it in a separate article in the Pi.dev-Toolbelt collection.

Handoff to a new session
Image description as text

Two equally wide context bars show two separate sessions. Above, session 1 contains the blue system prompt and its old history of requests, responses, a tool call and a tool output. The connecting arrow denotes summarising for the handoff. Below, session 2 begins only with a blue system prompt and a muted turquoise session handoff, followed by available context. There is neither a new user request nor an LLM response in the lower bar. The handoff transfers summarised information, not the entire old history. Both bars carry red Dumb Zone hatching in the right-hand portion of the available context. Its schematic starting point varies between the sessions. This personal guide is not a technical limit or memory reservation. Widths are schematic, not token measurements.

Compaction vs. handoff — a flamewar

There is no need to turn these approaches into rival camps. Compaction condenses the conversation so it can continue. A handoff shapes the context of a new session around the work ahead. Which approach fits depends on what you want to preserve and how the software handles it. Both can lose important details. Neither is better for every task.

In the next chapter

So far, we have looked at the information available to the model when it produces a response. How it selects the next token from the possible candidates is a different question. That is the subject of planned Chapter 08, “What do temperature and top-p do?”

Deeper into the Rabbit Hole

Videos in English

Texts

Sources

  1. Anthropic — Context windows.
  2. vLLM — Engine Arguments. Server configuration example.
  3. Anthropic — Effective context engineering for AI agents.
  4. Anthropic — Compaction overview.
  5. Mirco Blitz — pi-blitz-handoff — README v1.2.4. Published product version 1.2.4.
  6. Liu et al. — Lost in the Middle: How Language Models Use Long Contexts. Abstract read; the full paper was not reviewed for this article.