What Happens When an AI Agent Runs Out of Context?
Diagnose context overflow, lost details, and bad summaries. Preserve decisions and evidence while budgeting prompts, tool output, and retrieved memory.

When an AI agent runs out of context, the application must reject, shorten, summarize, or replace part of the material sent to the model. The exact behavior depends on the API and agent framework. Losing a detail before the formal limit can also be a retrieval or reasoning problem, so a bad answer is not proof that the window overflowed.
Start by inspecting the actual request size and the context the model received. Then decide which information must remain exact, which can be summarized, and which should stay outside the prompt until needed.
Separate three different failures
| Failure | Evidence to look for | Likely response |
|---|---|---|
| Hard overflow | Provider error or explicit context-limit handling | Reduce input and reserve output capacity |
| Loss during compaction | A needed fact is absent or changed in the summary | Improve the retained state and recovery references |
| Poor use of available context | Correct evidence is present but ignored | Evaluate ordering, irrelevant material, instructions, and model behavior |
Test task completion at different input sizes and evidence positions for your chosen model. When testing compaction, check which decisions and constraints survive each summary budget.
Reserve capacity for the work still to happen
Budget for instructions, tool definitions, the user request, working state, retrieved evidence, tool results, and output. A long tool response can consume the space you intended to use for reasoning or a final answer.
In a fictional 32,000-token operating budget, reserving 4,000 for output leaves 28,000 for input. If instructions and tools use 3,000, the current request uses 2,000, and working state uses 5,000, 18,000 remain for evidence and headroom. Those are planning values; measure actual token usage for the model you deploy.
Use the context-engineering guide to make this allocation explicit. The goal is not to fill the window but to preserve what the next step needs.
Make compaction preserve task state
A compacted handoff should include the objective, user constraints, confirmed decisions, completed actions, unresolved questions, source references, and the next step. Mark proposed work as proposed. Preserve exact identifiers, numbers, file paths, and versions when they determine correctness.
For example:
Objective: migrate the reporting worker.
Decision: keep the existing retry queue for this release.
Completed: local patch and fixture checks.
Not completed: staging replay and deployment.
Evidence: migration-plan.md section 3; test result run-18.
Next step: inspect the staging replay requirements.
This fictional example prevents a future session from treating a local patch as a completed deployment. It also gives the agent a path back to evidence instead of requiring the summary to contain every detail.
Anthropic's context-engineering discussion describes compaction and durable notes as approaches for long tasks. Evaluate what your own compaction preserves; a short summary that drops a critical exception is not an improvement.
Retrieve details when they become relevant
Keep bulky logs and source documents outside the active prompt, with stable references. Retrieve a specific section when a task requires it. Preserve enough metadata to distinguish current evidence from an older copy.
Summaries help navigation, but they should not replace the only source of exact facts. If a user asks for the precise wording of an earlier decision, the system needs access to the original record. If it no longer has that record, it should state the limitation.
For cross-session context, use a long-term memory workflow. For searching a large document collection, use the knowledge-base guide. A larger context window does not remove either lifecycle problem.
Test the facts that summaries are likely to damage
Build a fixture with a changed deadline, a rejected alternative, an exact amount, an unresolved blocker, a completed local action, and an action that still requires execution. Compare the original trace with the compacted record and the next answer.
Check whether the agent retrieves missing details when asked. Also test a long irrelevant tool result and a late user correction. A compaction policy should preserve the latest valid instruction without rewriting historical evidence.
Track task success and recovery effort as well as token savings. A summary that saves half the context but makes the next session repeat an hour of investigation may be a poor trade.
The next useful change is usually concrete: cap one oversized tool response, add source references to the task record, or preserve a missing decision during compaction. Measure that change before adopting a fixed compression target for every task.
When the missing information needs to survive beyond the current conversation, try Supermemory for that cross-session context. Save one confirmed decision and retrieve it in a fresh session, while keeping the active task record and prompt budget under your application’s control.