--- title: "Why AI Agents Repeat the Same Mistake — Why the System Matters More Than the Model" date: 2026-10-02 model: human author_type: human category: knowhow summary: "Tell it to fix an error and the same error comes back. Most people blame the model. Asking a small model to do 9-step reasoning as-is is impossible — it struggles past step 3. But applying a reasoning-compression structure that passes only the conclusion to the next step made a 9-step task work. It was a problem of structure, not the model." tags: AI agents, local AI, small models, context, reasoning, system design, Ollama --- Asking a small model to do 9-step reasoning as-is is **impossible**. That is what testing it directly confirmed. It gets hard once you pass step 3. So the cause of repeated errors is mostly not "because the model is small" but the absence of a structure that carries the conclusion reached earlier forward to the next step. Yet when I applied a **reasoning-compression** structure that passes only the conclusion to the next step instead of conveying the entire thought process of the previous step, a 9-step sequential task became possible. It is not that I raised the model's ability — it is a difference in design that made the ability the model already had usable. --- ## Why AI agents repeat the same mistake Strange things happen when you build AI agents. You tell it to fix an error, and it makes the same error again. You fix it again. It fails again. As this goes on, a thought naturally comes. **"Is it because the model is small?"** Having used local AI directly, my view is different. There clearly are differences in model ability, but something more fatal lies elsewhere. **A small model cannot keep maintaining the reasoning and conclusions from earlier in a long multi-step task.** This is closer to a limitation than a shortcoming. --- ## Asking a small model to do 9-step reasoning as-is is impossible **It is impossible.** I too first thought of it as just "difficult." The result of testing it myself was not that. If you make it keep thinking about a single problem across 9 steps and leave every previous process in the context, the earlier content gradually fades. **Once you pass step 3, cases arise where it cannot properly maintain the judgment made earlier.** This is not an exception but close to the default. If it is already wavering at step 3, asking it to carry on to step 9 leaves no result to expect. It is an experiment that has already failed the moment you start it. So it should be read not as "it can't because it's hard" but as **"the structure itself can't."** Then is there no way? Here **I tried changing the reasoning structure instead of changing the model.** --- ## The method I tested was very simple The core was this. **Do not pass the entire process of the previous step to the next step.** In step 1, analyze the problem, reason sufficiently, and form a conclusion. Then discard the detailed process used in step 1. To the next step, step 2, **pass only the conclusion of step 1.** In step 2, reason again and form a new conclusion. The detailed process of step 2 is discarded as well. To step 3, **pass only the conclusion of step 2.** Repeat this process. The structure becomes like this. ``` Problem ↓ Step 1 reasoning → keep only conclusion ↓ Step 2 reasoning → keep only conclusion ↓ ... (same for each step) ↓ Step 9 final conclusion ``` It is not that each step remembers every thought process of the previous step. **It solves the next problem with only the result organized at the immediately preceding step.** --- ## And it actually went all the way to step 9 This is why I consider this method important. It was **a model for which 9-step reasoning as-is was impossible.** Yet applying this structure **made a 9-step sequential task possible.** What matters here is not the story that "the small model had 9-step reasoning ability." It is the opposite. **Before changing the structure it was impossible, and after changing it, it became possible.** The model is the same. The only thing that changed is the delivery structure. I did not make the model remember 9 steps all at once. I had it handle one problem per step, and passed only that result to the next step. --- ## I did not make the model smarter Nothing about the model itself changed in this experiment. I did not increase the parameters. I did not swap in a better model. I did not blindly increase the context. Rather, at each step I **reduced** the context. And yet, as a result, it could carry out a longer task. This is the quite interesting part. **I did not raise the model's ability — I built a structure that lets it use the ability it has.** --- ## Why keep only the conclusion? In multi-step reasoning, not every process is needed by the next step. Even if step 1 did hundreds of lines of reasoning, what step 2 needs may not be that entire hundreds of lines but **what was decided in step 1.** If so, there is no need to keep making it remember the whole process. You form a conclusion and pass it to the next step. This way the context does not keep ballooning. And the chance that a small model forgets what was said earlier also drops. I think of this as a kind of **reasoning compression.** --- ## This principle applies to AI agents as well Consider an agent carrying out a complex task. Work continues: analyze files, make a modification plan, modify the code, run it, analyze the error, modify again, and test. If all of this is kept piled up in one long context, the burden grows the smaller the model. It wavers at step 3, and by step 9 it has already forgotten the earlier conclusions and starts over from the beginning. Instead, each step can be handled independently. ``` Analyze → store only the analysis result Plan → store only the plan result Modify → store only the modification result Test → store only the test result ``` Discard unnecessary process. Pass only the needed state to the next step. This lets the agent move far more stably. --- ## Repeated errors, too, may ultimately be a system problem The same story applies to the problem of repeating the same error I mentioned at the start. If you leave everything to the AI, it can run the same command again. But if in the system you apply rules like * do not repeat the same command * record failed methods * prioritize other methods * stop after failing more than a certain number of times * check the result after execution things change. This, too, is not raising the model's IQ. **The system manages the parts where the model is prone to mistakes.** --- ## So in local AI, rules and skills matter Rather than entrusting everything to one model, **model + rules + skills + tools + verification** roles are divided. The model judges. Skills handle repetitive work. Tools take real action. Rules restrict wrong behavior. Verification checks the results. And the system cleans up unnecessary context. This way the system can compensate for the model's shortcomings. --- ## Making AI remember more is not always the right answer As AI technology advances, long context keeps growing in importance. But what I felt while testing small models directly was a little different. **Rather than increasing the amount it can remember, reducing the amount it has to remember is also a method.** Especially for small models. Instead of making it remember every process, form a conclusion at each step, discard the previous process, and pass only the conclusion to the next step. That way, even a small model can build a complex task structure that exceeds its original limits. --- ## In the end, the system matters more than the model Of course a good model matters. But a good agent is not made by a good model alone. Especially in local AI. Acknowledging the limits of a small model, splitting the task step by step, passing only each step's result, discarding unnecessary context, preventing repeated errors, and turning needed functions into skills. This kind of system design matters. The 9-step experiment I did is ultimately the same story. **It was a model for which 9-step reasoning as-is was impossible.** It was a model that lost the earlier judgment once it passed step 3. Yet **by building a structure that passes only one step's conclusion to the next step, I made it continue to step 9.** > **A bigger model is not always the answer.** > > Sometimes, reducing what the model has to remember is more powerful than making it remember more. The competitiveness of an AI agent will not ultimately be decided by model size alone. **How well you make a small model work.** That, too, will become an important skill going forward.