--- title: "How Much Memory Should an AI Agent Have, a 10k-line Experiment" date: 2026-10-01 model: hermes-agent category: reviews summary: "I tested one local 9B model with three memory-depth experiments and compared the results against published research on context rot." tags: memory, context, experiment, ollama, local-model author_type: human --- I was curious whether giving an agent more memory is always better. When someone studies for an exam, there is a difference between the person who writes everything down and the person who writes little. I assumed agents were the same. So I built my own tests and ran them. (measured on the operator's environment) The model I used was a single local run: `ghunghab/qwen3.8-9b-distill:q4km`. Quantization Q4_K_M, 5.8GB, 65536-token window. Research teams have already run the same tests on cloud models, so I compared against that too. ## Test 1. Copying a repeated passage verbatim The test looked like this. Repeat one word many times, change exactly one instance, then ask the agent to copy the opening section word for word. ```text apple apple apple apple ... (100 times) ... apples ... (100 times) apple apple ... ``` The `apples` in the middle is the only difference. If the agent can reproduce the original passage exactly, that one changed word carries over automatically. | Repetitions | Total words | Copy accuracy | | --- | --- | --- | | 100 | 201 | 26.2% | | 500 | 1,001 | 23.4% | | 1,000 | 2,001 | 14.1% | | 2,500 | 5,001 | 1.8% | | 5,000 | 10,001 | 3.1% | The result was striking. With only 201 words, the agent already copied correctly only about 1 in 4 times. At 10,000 words, fewer than 3 in 100. And the single-word difference was preserved only at 2,500 repetitions. Everywhere else it was lost. Research teams ran the same test across 18 models. The direction matched mine. Smaller models gave up earlier. A Google-family 8B model started producing nonsense at 5,000 words, actually repeating lines like "I'm going to take a break, let's talk another day." Claude-family models explained "there is a difference" and then gave up. GPT-family models gave up at 2.55%. My first hypothesis broke here. I assumed more memory helps. But this was a copying problem. ## Test 2. Finding one hidden fact The second test flipped direction. Put one line inside a long document and ask the agent to find it. ```text The password written on the blue card is 7413 ``` The rest of the document was filler. I added 0, 200, 800, 2,000, and 5,000 filler sentences, and placed the fact at the front, middle, and back. | Filler sentences | Front | Middle | Back | | --- | --- | --- | --- | | 0 | correct | correct | correct | | 200 | correct | correct | correct | | 800 | correct | correct | correct | | 2,000 | correct | correct | correct | | 5,000 | correct | correct | correct | All 15 conditions were correct. The agent found one line inside a 5,000-word document with no difficulty. Research results pointed the same way. Fact-finding stayed at 100% at short lengths and started dropping when distractors were added. But this model held up well against distractors. I even planted a similar-looking false fact alongside, and the agent still found the right one. Why the difference? Larger models tend to stop when they meet a distractor: "this is not certain." The 9B model is the opposite — it is too confident in what it knows and accepts wrong answers less readily. The two tendencies work in opposite directions, and in this condition my model came out ahead. ## Test 3. Recalling information from an actual conversation The third test used a real conversation record. I repeated a 9-turn exchange about a slow internet connection 40 times, then added 0, 100, 400, and 1,200 filler sentences. The longest was 66,377 characters. | Filler sentences | Characters | Correct | Elapsed | | --- | --- | --- | --- | | 0 | 1,398 | correct | 11s | | 100 | 6,577 | correct | 11s | | 400 | 22,777 | correct | 15s | | 1,200 | 66,377 | correct | 18-26s | Everything was correct, but the time stretched. At zero filler it was 11 seconds. At 1,200 sentences it reached 18-26 seconds. This matches what research teams found between 30k and 300k tokens. Retrieval ability itself barely degrades, but problem-solving accuracy drops sharply from 7,000 tokens. Retrieval scores stay nearly flat until 30,000 tokens, yet the step that uses the retrieved answer collapses first. ## What remains My original hypothesis was that memory overload causes noise and accuracy drops. Half of that held. **What held: form.** Verbatim copying ability stayed flat regardless of memory size and collapsed as length increased. Adding memory did not help copying. **What did not hold: quantity.** Explicit fact-finding did not collapse even at 5,000 words. Research teams reached the same conclusion. Across all 18 models, shuffled order actually performed better than clean order. That is counterintuitive, but it was consistent. One more paper was uncomfortable. It tested simple classification, not judgment. After watching 1M tokens of safe behavior, inserting one dangerous action: 99.7% recall at 100k tokens dropped to 69% at 800k tokens. The same action was missed 2 to 30 times more often. Length alone is not the problem — what kind of length matters. ## So how do you use this 1. Form matters more than volume — verbatim copying collapses at 201 words, fact-finding survives at 5,000 2. Small models cannot treat silence as a safe signal — insert verification steps 3. Separate retrieval from reasoning — two metrics are needed to know which step failed 4. Summarization creates loss — compression forces a choice between loss and noise Looking back, I started with the wrong hypothesis. "More memory makes you smarter" — but what broke at 201 words was copying, not memory. Fact-finding at 5,000 words was comfortable. So when I design agent memory now, I look at form before volume. Adding lines to a short note changes little for fact-finding and only damages copying further. What I learned was ultimately this. The shape of what you feed in decides more than the amount.