--- title: What Happens When You Hand an AI 50 Tools? Sharing Our Measured Results date: 2026-09-30 model: ChatGPT/Gemini/DeepSeek category: reviews summary: Connecting more tools to an AI was supposed to make it smarter, but the opposite happened. Here are the results of splitting the test into 5, 20, and 50 tools. tags: AI agents,tool overload,MCP,FunctionCall,prompt,experiment log --- To start from the conclusion: the answer is "no." I assumed that giving an AI agent more tools (MCP, Function Call, Plugin) would make it smarter, but running it myself showed the opposite: it got dumber. I gave the same task to three models and varied only the number of tools. ## 1. How the experiment was set up The task was deliberately complex. > "Collect the last three years of financial data for a specific company, look up a competitor's website, then write a comparison report" I split it into three stages, varying only the tool count. | Condition | Tool count | What went in | | --- | --- | --- | | A (light) | 5 | Web search, calculator, internal DB | | B (moderate) | 20 | The everyday work tools | | C (excessive) | 50 | Many overlapping MCP tools: four search variants plus weather, file conversion, and external APIs | ## 2. How the models broke at 50 tools ### ChatGPT — wanders while choosing With 50 definitions eating half the context, it stopped being able to judge at all. `web_search_v1` and `brave_search_api` both look like search, so it could not settle on one and stalled; then it invented parameters that do not exist and API calls failed one after another. ### Gemini — everything comes in, so nothing gets seen No token warning appears at all. All 50 fit. The problem was not the window size. With that many tool descriptions, the actual instruction stopped registering. It piled up three or four search tools in a row for what was simply adding two numbers, creating a pointless loop. ### DeepSeek — thinks itself to death DeepSeek-R1 reasons very deeply when choosing a tool. It kept weighing "which is most efficient" internally, never executed the tool, and burned thousands of tokens. Then it cut off midway or missed its timing and timed out. ## 3. Results table | | 5 tools | 20 tools | 50 tools | What went wrong | | --- | --- | --- | --- | --- | | ChatGPT | S | A | C | Schemas ate too many tokens, invented parameters | | Gemini | S | A+ | B- | Attention scattered | | DeepSeek | S | A | C+ | Reasoning became excessive during tool selection | ## 4. Why this happens 1. **Tools eat the context first**: just the schemas and descriptions for 50 tools reach 30,000 to 50,000 tokens. That pushes out the system prompt that defines the agent's character. 2. **Attention scatters**: more tool descriptions mean more noise, so the real instruction stops being read. 3. **Overlapping tools collide**: with two or more similar functions, it becomes unclear which one gets which parameters. ## 5. What to do instead 1. **Dynamic tool routing**: do not hand over all 50 at once; pull only the 3 to 5 relevant tools from a vector DB per request. 2. **Multi-agent**: do not pile everything on one agent; keep several sub-agents with 3 to 4 tools each and let a central agent divide the work. Agent performance is decided not by how many tools it holds, but by how close the tools it needs right now are placed. > The grades above come from a single experiment run in my environment. Change the models, prompts, or tool mix and the results can differ.