The loop from the previous two articles — send messages, append the assistant message, run whatever tools it asked for, append the results, call again — has no notion of time. It does not care whether it is on its third turn or its fiftieth. That is exactly why it works for “Is it warmer in Paris or Dublin?” and starts quietly falling apart around step 40 of a real task.

Two things break, and they are not the same thing: the window runs out of room, and the model gets worse at using what is in it. Two families of fixes map onto them: compressing the window, and moving state out of it into an explicit plan the agent keeps re-reading.

The agent has exactly one memory

The Chat Completions and Messages APIs are stateless. Nothing about your task is remembered on the server between calls. Everything the model knows at step kk is what is inside the array you send at step kk.

That sounds obvious, but it has a sharp consequence: messages is not a record of the agent’s memory, it is the memory. Anything not in the array does not exist as far as the model is concerned. A file the agent wrote, a plan on disk, a row in a database — those are memory in the everyday sense, but they are inert until some code appends them back into the array as a message. Keeping this distinction straight explains both of the failures below and both of the fixes.

Failure 1: the window fills up

The hard limit is tokens. Each turn appends an assistant message and at least one tool result, and the next call re-sends everything. So the array grows linearly in steps, and the cumulative number of input tokens you send grows quadratically in steps.

Put numbers on it. Suppose each turn adds about 1,500 tokens (a short assistant message plus one tool result), and the task runs 50 turns — Manus reports that a typical task there averages around 50 tool calls 4. The final array holds about 50×1,500=75,00050 \times 1{,}500 = 75{,}000 tokens. The total you actually send is

∑k=1501500k=1500⋅50⋅512≈1.9M tokens\sum_{k=1}^{50} 1500k = 1500 \cdot \frac{50 \cdot 51}{2} \approx 1.9\text{M tokens}

about 25 times the final size. You do not pay for the context once; you pay for it on every call.

Prompt caching softens the bill a lot, because the prefix repeats exactly and vendors charge less for cached input. What caching does not do is shrink the window. All 75,000 tokens are still there on call 50, and the model still has to find things in them. Which brings us to the more interesting failure.

Failure 2: quality drops before the window fills

You might assume that a model with a 200,000-token window is equally reliable at token 200,000 and token 200. It is not.

Chroma’s July 2025 report, Context Rot, tested 18 frontier models on retrieval tasks and found accuracy degrading as input length grew — even on tasks as simple as pulling a fact out of text 3. Degradation got worse when the question and the target fact had low lexical overlap (so the model had to infer the connection rather than string-match it), and worse again with distractors present. The results also found that shuffling the haystack and destroying its narrative flow improved performance, which is a good hint that the failure is about attention, not comprehension. One caveat worth keeping: this is a vendor technical report, not peer-reviewed work, and its findings are tied to the specific models tested in mid-2025. The direction, though, matches what practitioners consistently see.

The intuitive explanation is that attention is a finite budget spread across the tokens present, and it is not uniform: the beginning and the end of a long context get more reliable use than the middle — the “lost in the middle” pattern 4. So a growing transcript does two things at once: the current goal sits further from the end, and more irrelevant material competes with it. This is why “wait for a bigger context window” is not a fix. Anthropic’s framing is the useful one: context engineering is about finding the smallest set of high-signal tokens that maximizes the chance of the outcome you want 1.

Fix 1: compress the window

The standard remedy is compaction: when the conversation approaches the limit, send the history to a model, get back a summary, and restart the array with that summary instead of the original messages 12.

The important property is that compaction is lossy by design. In production systems it cuts the token count by roughly 90–99%, and because the model decides what to keep, there is no guarantee that it keeps the same things across two runs of the same task 6. So what disappears first is verbatim detail: exact numbers, exact file contents, exact wording. High-level facts central to the task usually survive; an obscure figure from an appendix usually does not 2. Anthropic’s practical advice on tuning the summarization prompt is to maximize recall first — make sure the summary captures everything relevant — and only then tighten it for precision, because over-aggressive compaction drops subtle context whose importance only becomes obvious later 1.

There is a cheaper, narrower operation that is worth knowing before you reach for full compaction: tool-result clearing. Compaction rewrites the whole transcript; clearing walks the message list and surgically replaces old tool_result blocks with a short placeholder, keeping the tool_use record so the model still knows it made the call 2. That is safe precisely because tool output is usually re-fetchable — if the agent needs the file contents again, it can just read the file again. Note the shape of that argument: the reason clearing is cheap is that the information still exists somewhere else. That idea is about to become the main lever.

One trap, and it connects directly to the previous article. If you hand-roll compaction, you are cutting a hole in the middle of a message list that the API validates. If the cut falls between an assistant message carrying tool_calls and the tool messages answering it, you have orphaned tool-call ids and the next request is rejected. Vendor compaction features handle pairing across the summary boundary for you 2; if you implement your own, the boundary has to respect the grammar from the last article.

Fix 2: move state out of the window and read it back

Compression tries to keep the same information in a smaller space. The other fix does something better: it stops depending on the transcript to carry the state at all.

Structured note-taking means the agent writes notes to storage outside the context window — a NOTES.md, a todo.md, a scratch file — and pulls them back in later when it needs them 1. Because the notes live outside the array, they survive compaction and even survive the session ending. But remember the rule from the first section: they only affect the model when something appends them back in. Which is why the interesting part is when that happens.

Manus describes its agents maintaining a todo.md and rewriting it step by step, checking items off as they go. The stated reason is not bookkeeping. Rewriting the list recites the objective at the end of the context, where attention is strongest, pushing the global plan out of the lost-in-the-middle zone and reducing goal drift across a roughly 50-call run 4. The agent is spending tokens to steer its own attention, using nothing but natural language.

That mechanism is worth stating precisely, because it is easy to misread as “the agent remembers more.” It does not remember more. It re-states, on every turn, a compact version of the thing it must not lose — and the cost of that is tokens. The plan file is a lossy compression of the task’s state that is written by design rather than reconstructed later by a summarizer. That is the real split between the two fixes: summaries reconstruct state from history and hope nothing important was dropped; plan files create the compact state in the first place, so there is less to lose.

The cost is real too. Bookkeeping turns do not advance the task; one community write-up of the pattern reports that roughly a third of Manus’s actions went into updating the todo file, which is why later versions moved planning into a separate planner that directs executor sub-agents 7. Treat that number as second-hand and version-specific, but the trade-off it illustrates is genuine: the more you recite, the more of your budget goes to recitation.

Planning is an architectural choice, not a prompt

Once you accept that the plan should be an explicit, durable object rather than a feeling inside the model, you have a design question: who owns it, and when may it change?

At one end is the plain ReAct loop. The plan exists as prose inside assistant messages, revised silently each turn. The model can adapt to anything immediately, but the plan is also continuously overwritten, hard to inspect, and the first thing a compaction pass will blur.

At the other end is plan-and-execute: one component produces a structured list of steps, another carries them out. Plan-and-Act formalizes this as a Planner that emits high-level structured plans and an Executor that turns them into environment actions, with dynamic re-planning after each executor step 5. Separating the two reported 57.58% success on WebArena-Lite and 81.36% (text-only) on WebVoyager, and the authors’ explanation is a load-balancing one: a single model asked to hold a long-horizon strategy and produce the right low-level click on every step is carrying two jobs that interfere with each other 5.

The advantages beyond accuracy are the ones you get from any explicit artifact: the plan is auditable before execution, the failure point is identifiable, and the plan itself tells you what remains to do, so a context reset does not erase the task’s shape. The two standard pitfalls are also predictable. If there is no re-plan trigger, the executor marches through a plan that has already been invalidated by the environment. And a planner reasoning without tool results will happily write steps that depend on data it has never seen.

Where does the plain loop still win? Short tasks, and any task where the next step genuinely depends on what the last tool returned. Splitting planning from execution costs you the adaptability you get from deciding everything one step at a time. Anthropic’s general advice applies: start with the simplest thing that works, and add structure when measurement says you need it 8.

The rule that follows

Putting it together, the window at any single model call should look like this:

Rendering

Everything older than the last few turns has one of three destinations: folded into the current plan (if it is task state), cleared (if it is re-fetchable), or left on disk (if it might be needed later but is not needed now). Nothing stays in the array by default.

That is also where the boundaries are. Summaries lose verbatim detail and vary between runs, so anything you will need exactly — an ID, an exact number, a legal phrase — should be written somewhere durable and re-read, not trusted to a summary. Clearing only works on information you can fetch again. And a plan is a guess, so it needs a trigger for when it is allowed to change.

A useful first exercise: take the loop you already have and add only two things — a tool -message clearing rule that fires once the array passes N tokens, and a plan.md the agent rewrites at the start of every turn. Then run a 30-step task and watch which of the two actually keeps the agent on target. In most implementations, it is the second one.