Large tool outputs can overflow an agent's context and cause hallucinations or degraded responses.
2
Bigger context windows do not solve the problem because models tend to remember the beginning and end of very long contexts while losing the middle.
3
Agents can manage context by externalizing, selecting, compressing, and isolating information, then using pointers, memory, limits, and asynchronous tool calls where needed.
Summary
Elizabeth Fuentes Leone explains why agents degrade when tools return large amounts of data. A log-watching agent may add more logs to its context on every invocation, causing the context to grow while the agent tries to remember the whole session. She argues that increasing the context window is not enough because models can lose information in the middle of a long context. Her approach is to give the model only the information it needs at the time it needs it. She presents four strategies: externalize large data into storage, select relevant information, compress history through summarization, and isolate context between agents. The examples use Strands Agents, including conversation managers, short-term and long-term memory, graph memory, memory pointers, shared state across agent swarms, tool-call limits, and asynchronous handles for slow MCP tools. She closes with anti-patterns such as stuffing the context and passing every tool result to the model.
Large tool outputs make agent context grow until responses degrade
Leone describes an agent that supervises an application by retrieving its logs. Each retrieval adds more data to the context window, while the session also retains previous information. As the context grows, the engine can overflow, hallucinate, or produce degraded responses. The example shows why a log-watching agent cannot keep placing every result into its prompt and expect the model to remember the whole history accurately.
A larger context window does not prevent the model from losing the middle
Leone says that adding more tokens was once treated as the answer to context problems. She now considers that assumption false. Her explanation uses the attention curve: when the data becomes very large, the engine tends to remember the first and last parts of the context while forgetting information in the middle. The model therefore needs a better selection of information, even when the available window is large.
Context engineering gives the model only the information it needs
Leone defines context engineering as giving the model the information it needs when it needs it. An agent does not always need every item inside its context window, and excess information can make the context harmful. She also warns that a prompt can contain a prompt injection. Her goal is to control the contents of the context for better outcomes and lower token consumption.
Externalize, select, compress, and isolate are the four context strategies
Leone organizes her techniques around four actions. Externalize moves large data into persistent storage and leaves a memory pointer in the context. Select retrieves only relevant information. Compress reduces tokens by summarizing or compacting information. Isolate separates context across agents so that each agent receives only what it needs.
Strands Agents provides conversation managers and model choices
Using the open-source, model-agnostic Strands Agents framework, Leone shows three conversation-manager approaches. A sliding window keeps only the most recent messages. Summarization condenses older messages. A combined approach summarizes older messages while preserving recent ones. Her example keeps the last four messages and summarizes 50 percent of the older conversation. She also says the framework can use different model providers by adding the relevant model line.
Different memory types cover recent sessions, past information, and relationships
Leone separates memory into short-term, long-term, and relational forms. Short-term memory holds recent conversation history for the current session. Long-term memory can use a vector database so a later session can retrieve information from the past. Relational memory can use an entity graph, such as a graph that records how things connect. The three approaches can be combined when an application needs all of them.
Memory pointers keep large tool results out of the active context
For an application that produces many logs, Leone stores the logs instead of returning all of them to the model on every invocation. One tool saves the result in storage and returns an ID. That ID is kept in agent state as a memory pointer. Another tool can retrieve the logs when the agent actually needs them. The active context therefore contains a reference rather than the full large result.
Agents can share pointers without sharing every piece of context
Leone warns that sharing a whole context window across a multi-agent swarm adds irrelevant information to other agents. Instead, agents can share a pointer through invocation state. Each agent can use the ID to access the stored information when necessary, while the other agents avoid receiving the entire tool result.
Tool limits and asynchronous handles prevent stalled or endless agents
Leone describes agents that repeatedly invoke the same tool because they do not receive a clear response. A maximum tool count limits how many times a tool can run, with her example allowing a search tool three invocations. For slow MCP tools and external APIs, an asynchronous handle lets the agent start the job, save its status, and retrieve the answer during a later invocation instead of waiting for the request indefinitely.