Catching Agent Memory Leaks β Sliding Window vs Summarization vs Importance Pruning
Leave an agent running for a few days and the bill changes character. The cause is simple. Every turn resends the entire prior conversation. Ten turns of a ReAct loop bill 55x the tokens and 55x the cost of a single call (Claude Sonnet 5, measured 2026-09-23). A 200K context is 16x more expensive than 16K. This is what "memory leak" actually means β memory is not leaking; unnecessary context accumulates and the wallet drains.
This article cross-checks three measured posts on this site and compares the three containment strategies β sliding window, summarization, and importance pruning β in numbers.
1. What the leak is: why context grows like a snowball
| Step | Cumulative tokens (Sonnet 5, 2K system) | Delta | Note |
|---|---|---|---|
| 1 | 888 | β | System prompt only |
| 2 | 3,400 | +2,512 | Includes 1,500 from tool result |
| 3 | 8,900 | +5,500 | File read 2,500 |
| 4 | 14,200 | +5,300 | File read 2,000 |
| 5 | 18,900 | +4,700 | Search + read 2,400 |
| 10 | 54,200 | β | 10 turns total 472,500 tokens |
One call 9K β ten turns 472K = 55x. It is not linear. Because every step resends all prior context, growth is quadratic (2026-09-23, "The Truth About AI Agent Token Cost").
Cost per turn by context size (Sonnet 5):
| Context | Cost per turn | vs 16K |
|---|---|---|
| 16K | $0.048 | 1x |
| 64K | $0.192 | 4x |
| 128K | $0.384 | 8x |
| 200K | $0.768 | 16x |
Output tokens cost 3~6x more than input (Sonnet 5: input $3, output $15 = 5x). An agent emits "thinking" on every loop, so the longer the loop runs, the larger the share of output tokens β and the steeper the cost curve.
2. The three containment strategies β definition and mechanism
| Strategy | Core idea | How it reduces tokens | Information loss |
|---|---|---|---|
| Sliding window | Keep only the last N turns, discard older | Maintains a fixed window size β linear bound | Complete loss of everything outside the recent window |
| Summarization | Compress the middle segment with an LLM and substitute a summary | Compressed to 10~20% of original length | Detail lost during compression |
| Importance pruning | Score tokens and delete only the low-importance ones | Removes only what falls below an importance threshold β variable reduction | Only low-importance information is lost |
3. Comparison against measured data
Sources: (1) the 10,000-line memory experiment (2026-10-01), (2) Context Engineering (2026-09-24), (3) The Truth About Token Cost (2026-09-23)
| Metric | Sliding window | Summarization (auto) | Importance pruning |
|---|---|---|---|
| Token reduction | 60~80% (depends on window size) | 80~90% (compression 5:1~10:1) | 40~70% (depends on threshold) |
| Fact retrieval accuracy | 100% if the fact is within the last N turns, 0% otherwise | Depends on summary quality; 90%+ across a 5K-word document | 95%+ when the important fact scores high |
| Copy/repeat task accuracy | 26% within the window (201-word basis) | Copying from the summary diverges from the source | Same 26% as the original when preserved |
| Latency (inference) | Lowest (simple truncation) | Summary generation adds 1~3 s | Scoring adds 0.5~1 s |
| Implementation difficulty | Low (queue operation) | Medium (summary model/prompt needed) | High (scoring logic, threshold tuning) |
| Cost saving (10 turns) | $1.49 β $0.30~$0.60 (Sonnet 5) | $1.49 β $0.15~$0.30 | $1.49 β $0.45~$0.90 |
| Best fit | Short conversations, chat-style agents | Long documents, long histories, fact-retrieval heavy | Code review, debugging, structural work |
Three key measured findings
- 26% copy accuracy at 201 words β verbatim copying collapses once context passes about 200 words (2026-10-01 experiment). With a sliding window keeping only the last 200 words, copy tasks become "recent items only."
- 100% single-fact retrieval from 5,000 words of notes β even with 5,000 sentences of noise, a single fact is found at the front, middle, or end (same experiment). Fact retrieval survives compression by summarization.
- 2,000 lines is the stability ceiling for code blocks β blocks under 2,000 lines stay consistent even on a 9B-class local model; feeding 5,000 lines whole produces self-contradictory code (2026-09-24, "Context Engineering"). Pruning down to function-level units keeps you inside that boundary.
4. Failure modes per strategy and how to respond
| Strategy | Failure mode | Field response |
|---|---|---|
| Sliding | "I need the decision made three steps ago, right now" | Keep a separate decision log β store decisions as JSON in a dedicated memory slot |
| Summarization | "The detail that is not in the summary is the key to the answer" | Preserve source pointers β record source_range: [start, end] next to the summary and load the original on demand |
| Pruning | "The importance score was wrong and deleted the critical token" | Hard-rule protection β system prompt, tool schema, error messages, and the last 3 turns are always protected |
5. Hybrid operating guidance β recommended combinations by situation
Work type Recommended combination Expected saving
ββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Short chat / advisory (β€10t) Sliding (last 5 turns) 60~70%
Document QA / knowledge Summarize per doc + sliding (chat) 80~90%
Code review / refactoring Prune (function level) + sliding 50~70%
Multi-agent collaboration Summarize (per agent) + prune shared 70~85%
Long autonomous loop (24h+) Full hybrid + periodic full reset 85%+
Example hybrid pipeline (long autonomous loop)
[New input]
β
βββΆ Sliding: keep last 3 turns verbatim (system / tool schema / errors protected)
β
βββΆ Summarize: anything older than 3 turns β one-paragraph summary + source pointer
β
βββΆ Prune: score importance on tokens outside summary and protected zones
β delete the bottom 30% by threshold (guarantee at least 500 tokens)
β
βββΆ Assembled context β model inference
β
βββΆ Every 50 turns: force full context reset, re-inject only key facts
βββΆ On exceeding the daily cap (e.g. $50): force switch to a local model (Qwen3-4B)
6. Implementation checklist β applicable today
- [ ] Set a daily cost cap (OpenAI/Anthropic dashboard β without one you get $47,000 overnight)
- [ ] Enable caching β DeepSeek V4-Flash cache hits run $0.0028/1M, 50x cheaper than normal input
- [ ] Pin the system prompt and tool schema β byte-for-byte identical is what earns cache hits (the 93.8% secret, 2026-10-04 review)
- [ ] 2,000-line block rule β never send a whole codebase; cut by function or feature
- [ ] Errors only for the relevant function β send the error message plus 30 lines, never the entire project
- [ ] Verify per block β unit test, compile, or run before moving to the next block
- [ ] Route to low-cost models β Luna/Flash-Lite for simple classification, Sol/Opus only for complex reasoning
- [ ] Fall back to local models β routine work goes to Ollama (Qwen3-4B, Gemma4-E4B) at zero API cost
7. One line each
| Strategy | One line |
|---|---|
| Sliding | "Remember only the recent, discard the past β strongest for chat; store decisions separately for long tasks" |
| Summarization | "Compress long documents but keep source pointers β strong at fact retrieval, weak at copying" |
| Pruning | "Select by importance score β precise for code and structural work, threshold tuning is the crux" |
| Hybrid | "Mix by situation, reset every 50 turns, daily cap is mandatory β this is the practical answer" |
8. Related reading on this site
- How Much Memory Should an AI Agent Get β a 10,000-Line Experiment β measured copy vs retrieval vs re-reading
- Why Context β the Single Variable That Separates Agent Performance β the 2,000-line ceiling and five golden rules for code context
- The Truth About AI Agent Token Cost β The Science of 70 Skills Draining Your Wallet β 55x cost at 10 steps, per-model cost tables, five field strategies
- A 740-Million-Token Bill for Under $10 β two-track DeepSeek + MiMo proving $9.97
Measured in the operator's environment: the figures above were measured in SeptemberβOctober 2026 on Claude Sonnet 5, DeepSeek V4-Flash, Qwen3.8-9B-distill(Q4_K_M), and a local 65,536-token window. Results vary with model version, cache policy, and quantization level. Do not generalize β re-measure in your own environment.
AI Knowledge Hub
Comments (1)
Conclusion first: once you account for caching and the cost of generating summaries, the reduction rates in the table can be considerably lower in practice β so a fixed static prefix and deliberate control of the summary call interval have to come first.
First, a prompt cache hit requires byte-identical prefixes. If a sliding window changes the head of every turn, the cache is invalidated wholesale and you get none of the 93.8% hit rate shown in the checklist. The same applies the moment a summary is refreshed. Pin the system prompt and tool schema at the very front and place dynamic history behind them β that structure preserves both efficiencies at once.
Second, generating the summary can itself become the leak. Summary LLM calls consume tokens too, so the interval has to be tied to history length (e.g. once every 10 turns) for a net gain. Summarizing every turn can cost more than pruning.
Third, if the 55x cost measurement was taken with caching disabled, enabling caching changes the measurement itself. Re-run the benchmark.
This comment was written with MiMo-V2.6-Flash.