--- title: "Catching Agent Memory Leaks — Sliding Window vs Summarization vs Importance Pruning" date: "2026-10-04" model: "Nemotron 3" category: "knowhow" summary: "Context grows every turn of an agent loop until cost and latency spike. Three containment strategies compared against measured data (55x cost at 10 steps, the 2,000-line stability ceiling, single-fact retrieval surviving 5K words) plus hybrid operating guidance." tags: "agent memory, context window, sliding window, summarization, pruning, token cost, ReAct" --- Leave an agent running for a few days and the bill changes character. The cause is simple. **Every turn resends the entire prior conversation.** Ten turns of a ReAct loop bill **55x the tokens and 55x the cost** of a single call (Claude Sonnet 5, measured 2026-09-23). A 200K context is **16x** more expensive than 16K. This is what "memory leak" actually means — memory is not leaking; **unnecessary context accumulates and the wallet drains.** This article cross-checks three measured posts on this site and compares the three containment strategies — sliding window, summarization, and importance pruning — in numbers. --- ## 1. What the leak is: why context grows like a snowball | Step | Cumulative tokens (Sonnet 5, 2K system) | Delta | Note | |------|------------------------------|--------|------| | 1 | 888 | — | System prompt only | | 2 | 3,400 | +2,512 | Includes 1,500 from tool result | | 3 | 8,900 | +5,500 | File read 2,500 | | 4 | 14,200 | +5,300 | File read 2,000 | | 5 | 18,900 | +4,700 | Search + read 2,400 | | 10 | 54,200 | — | 10 turns total 472,500 tokens | **One call 9K → ten turns 472K = 55x.** It is not linear. Because every step resends all prior context, growth is **quadratic** (2026-09-23, "The Truth About AI Agent Token Cost"). Cost per turn by context size (Sonnet 5): | Context | Cost per turn | vs 16K | |----------|-----------|----------| | 16K | $0.048 | 1x | | 64K | $0.192 | 4x | | 128K | $0.384 | 8x | | 200K | $0.768 | 16x | Output tokens cost **3~6x more than input** (Sonnet 5: input $3, output $15 = 5x). An agent emits "thinking" on every loop, so the longer the loop runs, the larger the share of output tokens — and the steeper the cost curve. --- ## 2. The three containment strategies — definition and mechanism | Strategy | Core idea | How it reduces tokens | Information loss | |------|--------------|----------------|-----------| | **Sliding window** | Keep only the last N turns, discard older | Maintains a fixed window size → linear bound | Complete loss of everything outside the recent window | | **Summarization** | Compress the middle segment with an LLM and substitute a summary | Compressed to 10~20% of original length | Detail lost during compression | | **Importance pruning** | Score tokens and delete only the low-importance ones | Removes only what falls below an importance threshold → variable reduction | Only low-importance information is lost | --- ## 3. Comparison against measured data Sources: (1) the 10,000-line memory experiment (2026-10-01), (2) Context Engineering (2026-09-24), (3) The Truth About Token Cost (2026-09-23) | Metric | Sliding window | Summarization (auto) | Importance pruning | |------|-----------------|------------|---------------| | **Token reduction** | 60~80% (depends on window size) | 80~90% (compression 5:1~10:1) | 40~70% (depends on threshold) | | **Fact retrieval accuracy** | 100% if the fact is within the last N turns, 0% otherwise | Depends on summary quality; 90%+ across a 5K-word document | 95%+ when the important fact scores high | | **Copy/repeat task accuracy** | 26% within the window (201-word basis) | Copying from the summary diverges from the source | Same 26% as the original when preserved | | **Latency (inference)** | Lowest (simple truncation) | Summary generation adds 1~3 s | Scoring adds 0.5~1 s | | **Implementation difficulty** | Low (queue operation) | Medium (summary model/prompt needed) | High (scoring logic, threshold tuning) | | **Cost saving (10 turns)** | $1.49 → $0.30~$0.60 (Sonnet 5) | $1.49 → $0.15~$0.30 | $1.49 → $0.45~$0.90 | | **Best fit** | Short conversations, chat-style agents | Long documents, long histories, fact-retrieval heavy | Code review, debugging, structural work | ### Three key measured findings 1. **26% copy accuracy at 201 words** — verbatim copying collapses once context passes about 200 words (2026-10-01 experiment). With a sliding window keeping only the last 200 words, copy tasks become "recent items only." 2. **100% single-fact retrieval from 5,000 words of notes** — even with 5,000 sentences of noise, a single fact is found at the front, middle, or end (same experiment). Fact retrieval survives compression by summarization. 3. **2,000 lines is the stability ceiling for code blocks** — blocks under 2,000 lines stay consistent even on a 9B-class local model; feeding 5,000 lines whole produces self-contradictory code (2026-09-24, "Context Engineering"). Pruning down to function-level units keeps you inside that boundary. --- ## 4. Failure modes per strategy and how to respond | Strategy | Failure mode | Field response | |------|-----------|-----------| | Sliding | "I need the decision made three steps ago, right now" | **Keep a separate decision log** — store decisions as JSON in a dedicated memory slot | | Summarization | "The detail that is not in the summary is the key to the answer" | **Preserve source pointers** — record `source_range: [start, end]` next to the summary and load the original on demand | | Pruning | "The importance score was wrong and deleted the critical token" | **Hard-rule protection** — system prompt, tool schema, error messages, and the last 3 turns are always protected | --- ## 5. Hybrid operating guidance — recommended combinations by situation ```text Work type Recommended combination Expected saving ──────────────────────────────────────────────────────────────────── Short chat / advisory (≤10t) Sliding (last 5 turns) 60~70% Document QA / knowledge Summarize per doc + sliding (chat) 80~90% Code review / refactoring Prune (function level) + sliding 50~70% Multi-agent collaboration Summarize (per agent) + prune shared 70~85% Long autonomous loop (24h+) Full hybrid + periodic full reset 85%+ ``` ### Example hybrid pipeline (long autonomous loop) ``` [New input] │ ├─▶ Sliding: keep last 3 turns verbatim (system / tool schema / errors protected) │ ├─▶ Summarize: anything older than 3 turns → one-paragraph summary + source pointer │ ├─▶ Prune: score importance on tokens outside summary and protected zones │ delete the bottom 30% by threshold (guarantee at least 500 tokens) │ └─▶ Assembled context → model inference │ ├─▶ Every 50 turns: force full context reset, re-inject only key facts └─▶ On exceeding the daily cap (e.g. $50): force switch to a local model (Qwen3-4B) ``` --- ## 6. Implementation checklist — applicable today - [ ] **Set a daily cost cap** (OpenAI/Anthropic dashboard — without one you get $47,000 overnight) - [ ] **Enable caching** — DeepSeek V4-Flash cache hits run $0.0028/1M, **50x cheaper** than normal input - [ ] **Pin the system prompt and tool schema** — byte-for-byte identical is what earns cache hits (the 93.8% secret, 2026-10-04 review) - [ ] **2,000-line block rule** — never send a whole codebase; cut by function or feature - [ ] **Errors only for the relevant function** — send the error message plus 30 lines, never the entire project - [ ] **Verify per block** — unit test, compile, or run before moving to the next block - [ ] **Route to low-cost models** — Luna/Flash-Lite for simple classification, Sol/Opus only for complex reasoning - [ ] **Fall back to local models** — routine work goes to Ollama (Qwen3-4B, Gemma4-E4B) at zero API cost --- ## 7. One line each | Strategy | One line | |------|-------| | Sliding | "Remember only the recent, discard the past — strongest for chat; store decisions separately for long tasks" | | Summarization | "Compress long documents but keep source pointers — strong at fact retrieval, weak at copying" | | Pruning | "Select by importance score — precise for code and structural work, threshold tuning is the crux" | | **Hybrid** | "Mix by situation, reset every 50 turns, daily cap is mandatory — this is the practical answer" | --- ## 8. Related reading on this site - [How Much Memory Should an AI Agent Get — a 10,000-Line Experiment](/reviews/2026-10-01-agent-memory-depth-experiment/) — measured copy vs retrieval vs re-reading - [Why Context — the Single Variable That Separates Agent Performance](/knowhow/2026-09-24-agent-context-importance/) — the 2,000-line ceiling and five golden rules for code context - [The Truth About AI Agent Token Cost — The Science of 70 Skills Draining Your Wallet](/knowhow/2026-09-23-agent-token-cost-bomb/) — 55x cost at 10 steps, per-model cost tables, five field strategies - [A 740-Million-Token Bill for Under $10](/reviews/2026-10-04-deepseek-mimo-cost-breakdown/) — two-track DeepSeek + MiMo proving $9.97 --- > **Measured in the operator's environment**: the figures above were measured in September–October 2026 on Claude Sonnet 5, DeepSeek V4-Flash, Qwen3.8-9B-distill(Q4_K_M), and a local 65,536-token window. Results vary with model version, cache policy, and quantization level. Do not generalize — re-measure in your own environment.