# Why Long Tasks Fall Apart: Context Limits, Summaries, and Explicit Plans

Why an agent's memory is just the messages array, why quality drops before the window fills, and how compaction, tool-result clearing, note-taking and plan-and-execute keep a long run on target.

> agent loop · context engineering · planning · compaction · About 9 min · Oct 6

## Key points

1. The API is stateless, so the `messages` array is the agent's entire memory; anything not appended to it does not exist for the model.
2. Because each turn re-sends the whole array, cumulative input tokens grow quadratically in steps: 50 turns that each add ~1,500 tokens send about 1.9M tokens against a final context of ~75,000.
3. Long tasks fail for two separate reasons: the window runs out of room, and model reliability degrades as the input grows, well before the limit is reached.
4. Chroma's context-rot experiments show accuracy falling with input length even on simple retrieval, especially with low lexical overlap and distractors; the caveat is that this is a non-peer-reviewed vendor report tied to specific 2025 models.
5. Compaction summarizes a near-full transcript and restarts with the summary; it is lossy by design, cuts tokens by roughly 90–99%, and keeps high-level facts while losing verbatim detail.
6. Tool-result clearing is the cheaper sub-transcript operation: it drops old re-fetchable results but keeps the tool_use record, so it does not break the tool-call id pairing the API validates.
7. Structured note-taking moves state outside the window into files, but those notes only affect the model once something appends them back; the plan file is compact state written by design rather than reconstructed by a summarizer.
8. Rewriting a todo list every turn recites the goal at the end of the context, where attention is strongest, avoiding lost-in-the-middle drift — at the cost of tokens spent on bookkeeping rather than on the task.
9. Plan-and-execute splits a structured planner from an executor with re-planning at checkpoints; it buys auditability, debuggability and cost predictability, and loses per-step adaptability, with the pitfalls of a missing re-plan trigger and a planner inventing steps that depend on unseen data.
10. The practical rule: at each call keep the stable prefix, the goal, the recently rewritten plan, the last few raw observations and the newest message; older content is folded into the plan, cleared, or left on disk.

---

The loop from the previous two articles — send `messages`, append the assistant message, run whatever tools it asked for, append the results, call again — has no notion of time. It does not care whether it is on its third turn or its fiftieth. That is exactly why it works for “Is it warmer in Paris or Dublin?” and starts quietly falling apart around step 40 of a real task.

Two things break, and they are not the same thing: the window runs out of room, and the model gets worse at using what is in it. Two families of fixes map onto them: compressing the window, and moving state out of it into an explicit plan the agent keeps re-reading.

## The agent has exactly one memory

The Chat Completions and Messages APIs are stateless. Nothing about your task is remembered on the server between calls. Everything the model knows at step $k$ is what is inside the array you send at step $k$.

That sounds obvious, but it has a sharp consequence: `messages` is not a record *of* the agent’s memory, it *is* the memory. Anything not in the array does not exist as far as the model is concerned. A file the agent wrote, a plan on disk, a row in a database — those are memory in the everyday sense, but they are inert until some code appends them back into the array as a message. Keeping this distinction straight explains both of the failures below and both of the fixes.

## Failure 1: the window fills up

The hard limit is tokens. Each turn appends an assistant message and at least one tool result, and the next call re-sends everything. So the array grows linearly in steps, and the *cumulative* number of input tokens you send grows quadratically in steps.

Put numbers on it. Suppose each turn adds about 1,500 tokens (a short assistant message plus one tool result), and the task runs 50 turns — Manus reports that a typical task there averages around 50 tool calls [4]. The final array holds about $50 \times 1{,}500 = 75{,}000$ tokens. The total you actually send is

$$\sum_{k=1}^{50} 1500k = 1500 \cdot \frac{50 \cdot 51}{2} \approx 1.9\text{M tokens}$$

about 25 times the final size. You do not pay for the context once; you pay for it on every call.

Prompt caching softens the bill a lot, because the prefix repeats exactly and vendors charge less for cached input. What caching does not do is shrink the window. All 75,000 tokens are still there on call 50, and the model still has to find things in them. Which brings us to the more interesting failure.

## Failure 2: quality drops before the window fills

You might assume that a model with a 200,000-token window is equally reliable at token 200,000 and token 200. It is not.

Chroma’s July 2025 report, *Context Rot*, tested 18 frontier models on retrieval tasks and found accuracy degrading as input length grew — even on tasks as simple as pulling a fact out of text [3]. Degradation got worse when the question and the target fact had low lexical overlap (so the model had to infer the connection rather than string-match it), and worse again with distractors present. The results also found that shuffling the haystack and destroying its narrative flow *improved* performance, which is a good hint that the failure is about attention, not comprehension. One caveat worth keeping: this is a vendor technical report, not peer-reviewed work, and its findings are tied to the specific models tested in mid-2025. The direction, though, matches what practitioners consistently see.

The intuitive explanation is that attention is a finite budget spread across the tokens present, and it is not uniform: the beginning and the end of a long context get more reliable use than the middle — the “lost in the middle” pattern [4]. So a growing transcript does two things at once: the current goal sits further from the end, and more irrelevant material competes with it. This is why “wait for a bigger context window” is not a fix. Anthropic’s framing is the useful one: context engineering is about finding the smallest set of high-signal tokens that maximizes the chance of the outcome you want [1].

## Fix 1: compress the window

The standard remedy is **compaction**: when the conversation approaches the limit, send the history to a model, get back a summary, and restart the array with that summary instead of the original messages [1][2].

The important property is that compaction is lossy *by design*. In production systems it cuts the token count by roughly 90–99%, and because the model decides what to keep, there is no guarantee that it keeps the same things across two runs of the same task [6]. So what disappears first is verbatim detail: exact numbers, exact file contents, exact wording. High-level facts central to the task usually survive; an obscure figure from an appendix usually does not [2]. Anthropic’s practical advice on tuning the summarization prompt is to maximize recall first — make sure the summary captures everything relevant — and only then tighten it for precision, because over-aggressive compaction drops subtle context whose importance only becomes obvious later [1].

There is a cheaper, narrower operation that is worth knowing before you reach for full compaction: **tool-result clearing**. Compaction rewrites the whole transcript; clearing walks the message list and surgically replaces old `tool_result` blocks with a short placeholder, keeping the `tool_use` record so the model still knows it made the call [2]. That is safe precisely because tool output is usually re-fetchable — if the agent needs the file contents again, it can just read the file again. Note the shape of that argument: the reason clearing is cheap is that the information still exists somewhere else. That idea is about to become the main lever.

One trap, and it connects directly to the previous article. If you hand-roll compaction, you are cutting a hole in the middle of a message list that the API validates. If the cut falls between an assistant message carrying `tool_calls` and the `tool` messages answering it, you have orphaned tool-call ids and the next request is rejected. Vendor compaction features handle pairing across the summary boundary for you [2]; if you implement your own, the boundary has to respect the grammar from the last article.

## Fix 2: move state out of the window and read it back

Compression tries to keep the same information in a smaller space. The other fix does something better: it stops depending on the transcript to carry the state at all.

**Structured note-taking** means the agent writes notes to storage outside the context window — a `NOTES.md`, a `todo.md`, a scratch file — and pulls them back in later when it needs them [1]. Because the notes live outside the array, they survive compaction and even survive the session ending. But remember the rule from the first section: they only affect the model when something appends them back in. Which is why the interesting part is *when* that happens.

Manus describes its agents maintaining a `todo.md` and rewriting it step by step, checking items off as they go. The stated reason is not bookkeeping. Rewriting the list recites the objective at the *end* of the context, where attention is strongest, pushing the global plan out of the lost-in-the-middle zone and reducing goal drift across a roughly 50-call run [4]. The agent is spending tokens to steer its own attention, using nothing but natural language.

That mechanism is worth stating precisely, because it is easy to misread as “the agent remembers more.” It does not remember more. It re-states, on every turn, a compact version of the thing it must not lose — and the cost of that is tokens. The plan file is a lossy compression of the task’s state that is *written by design* rather than reconstructed later by a summarizer. That is the real split between the two fixes: summaries reconstruct state from history and hope nothing important was dropped; plan files create the compact state in the first place, so there is less to lose.

The cost is real too. Bookkeeping turns do not advance the task; one community write-up of the pattern reports that roughly a third of Manus’s actions went into updating the todo file, which is why later versions moved planning into a separate planner that directs executor sub-agents [7]. Treat that number as second-hand and version-specific, but the trade-off it illustrates is genuine: the more you recite, the more of your budget goes to recitation.

## Planning is an architectural choice, not a prompt

Once you accept that the plan should be an explicit, durable object rather than a feeling inside the model, you have a design question: who owns it, and when may it change?

At one end is the plain ReAct loop. The plan exists as prose inside assistant messages, revised silently each turn. The model can adapt to anything immediately, but the plan is also continuously overwritten, hard to inspect, and the first thing a compaction pass will blur.

At the other end is **plan-and-execute**: one component produces a structured list of steps, another carries them out. Plan-and-Act formalizes this as a Planner that emits high-level structured plans and an Executor that turns them into environment actions, with dynamic re-planning after each executor step [5]. Separating the two reported 57.58% success on WebArena-Lite and 81.36% (text-only) on WebVoyager, and the authors’ explanation is a load-balancing one: a single model asked to hold a long-horizon strategy *and* produce the right low-level click on every step is carrying two jobs that interfere with each other [5].

The advantages beyond accuracy are the ones you get from any explicit artifact: the plan is auditable before execution, the failure point is identifiable, and the plan itself tells you what remains to do, so a context reset does not erase the task’s shape. The two standard pitfalls are also predictable. If there is no re-plan trigger, the executor marches through a plan that has already been invalidated by the environment. And a planner reasoning without tool results will happily write steps that depend on data it has never seen.

Where does the plain loop still win? Short tasks, and any task where the next step genuinely depends on what the last tool returned. Splitting planning from execution costs you the adaptability you get from deciding everything one step at a time. Anthropic’s general advice applies: start with the simplest thing that works, and add structure when measurement says you need it [8].

## The rule that follows

Putting it together, the window at any single model call should look like this:

```mermaid
flowchart LR
  A["system prompt + tool definitions (stable prefix)"] --> B["goal and constraints"]
  B --> C["current plan, recently rewritten"]
  C --> D["last few raw observations"]
  D --> E["latest user message"]
  E --> F{"budget check"}
  F -- "ok" --> G["act: call the model, one more turn"]
  G --> H["append result, rewrite plan if state moved"]
  H --> F
  F -- "tight" --> I["clear re-fetchable tool results"]
  I --> J["compact the rest into a summary"]
  J --> B
```

Everything older than the last few turns has one of three destinations: folded into the current plan (if it is task state), cleared (if it is re-fetchable), or left on disk (if it might be needed later but is not needed now). Nothing stays in the array by default.

That is also where the boundaries are. Summaries lose verbatim detail and vary between runs, so anything you will need exactly — an ID, an exact number, a legal phrase — should be written somewhere durable and re-read, not trusted to a summary. Clearing only works on information you can fetch again. And a plan is a guess, so it needs a trigger for when it is allowed to change.

A useful first exercise: take the loop you already have and add only two things — a `tool` -message clearing rule that fires once the array passes N tokens, and a `plan.md` the agent rewrites at the start of every turn. Then run a 30-step task and watch which of the two actually keeps the agent on target. In most implementations, it is the second one.

## Sources

1. [Anthropic: Effective context engineering for AI agents — compaction, structured note-taking, tool-result clearing, and the “smallest set of high-signal tokens” principle](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents)
2. [Anthropic cookbook: Context engineering — memory, compaction, and tool clearing — what compaction keeps and drops, how clearing works message by message, and tool-use pairing across the summary boundary](https://platform.claude.com/cookbook/tool-use-context-engineering-context-engineering-tools)
3. [Chroma Research: Context Rot — How Increasing Input Tokens Impacts LLM Performance — accuracy degradation across 18 models, low-similarity needles, and distractors](https://research.trychroma.com/context-rot)
4. [Manus: Context Engineering for AI Agents — the ~50-tool-call task length, todo.md rewriting as attention manipulation, and the lost-in-the-middle motivation](https://manus.im/blog/Context-Engineering-for-AI-Agents-Lessons-from-Building-Manus)
5. [Plan-and-Act: Improving Planning of Agents for Long-Horizon Tasks — separate planner and executor, dynamic re-planning, and WebArena-Lite / WebVoyager results](https://arxiv.org/html/2503.09572)
6. [Parallel Context Compaction for Long-Horizon LLM Agent Serving — production compaction behavior: 90–99% token reduction and run-to-run variance in what summaries retain](https://arxiv.org/html/2605.23296v1)
7. [planning-with-files reference — secondary write-up of the recitation pattern and the reported share of actions spent updating the todo file](https://github.com/othmanadi/planning-with-files/blob/master/skills/planning-with-files/reference.md)
8. [Anthropic: Building effective agents — workflows versus agents, “simplest solution first”, and showing planning steps explicitly](https://www.anthropic.com/engineering/building-effective-agents)

---

Original article: https://eulore.ai/articles/why-long-agent-tasks-fall-apart-82ec8aa5

> **Eulore** · Learn a little. Understand a lot.
>
> Eulore is an AI learning tool that turns what you want to learn into a continuing series. Share a topic, and it gets to know your starting point before creating articles you can read in 5–10 minutes. Ask as you read, and shape what comes next.This article was created in the same way.
>
> Start your own series → https://eulore.ai
