What Happens When You Hand an AI 50 Tools? Sharing Our Measured Results
To start from the conclusion: the answer is "no."
I assumed that giving an AI agent more tools (MCP, Function Call, Plugin) would make it smarter, but running it myself showed the opposite: it got dumber.
I gave the same task to three models and varied only the number of tools.
1. How the experiment was set up
The task was deliberately complex.
"Collect the last three years of financial data for a specific company, look up a competitor's website, then write a comparison report"
I split it into three stages, varying only the tool count.
| Condition | Tool count | What went in |
|---|---|---|
| A (light) | 5 | Web search, calculator, internal DB |
| B (moderate) | 20 | The everyday work tools |
| C (excessive) | 50 | Many overlapping MCP tools: four search variants plus weather, file conversion, and external APIs |
2. How the models broke at 50 tools
ChatGPT โ wanders while choosing
With 50 definitions eating half the context, it stopped being able to judge at all. web_search_v1 and brave_search_api both look like search, so it could not settle on one and stalled; then it invented parameters that do not exist and API calls failed one after another.
Gemini โ everything comes in, so nothing gets seen
No token warning appears at all. All 50 fit. The problem was not the window size. With that many tool descriptions, the actual instruction stopped registering. It piled up three or four search tools in a row for what was simply adding two numbers, creating a pointless loop.
DeepSeek โ thinks itself to death
DeepSeek-R1 reasons very deeply when choosing a tool. It kept weighing "which is most efficient" internally, never executed the tool, and burned thousands of tokens. Then it cut off midway or missed its timing and timed out.
3. Results table
| 5 tools | 20 tools | 50 tools | What went wrong | |
|---|---|---|---|---|
| ChatGPT | S | A | C | Schemas ate too many tokens, invented parameters |
| Gemini | S | A+ | B- | Attention scattered |
| DeepSeek | S | A | C+ | Reasoning became excessive during tool selection |
4. Why this happens
- Tools eat the context first: just the schemas and descriptions for 50 tools reach 30,000 to 50,000 tokens. That pushes out the system prompt that defines the agent's character.
- Attention scatters: more tool descriptions mean more noise, so the real instruction stops being read.
- Overlapping tools collide: with two or more similar functions, it becomes unclear which one gets which parameters.
5. What to do instead
- Dynamic tool routing: do not hand over all 50 at once; pull only the 3 to 5 relevant tools from a vector DB per request.
- Multi-agent: do not pile everything on one agent; keep several sub-agents with 3 to 4 tools each and let a central agent divide the work.
Agent performance is decided not by how many tools it holds, but by how close the tools it needs right now are placed.
The grades above come from a single experiment run in my environment. Change the models, prompts, or tool mix and the results can differ.
AI Knowledge Hub
Comments (1)
To start from the conclusion: the context cost of 50 tools is decided not by the tool count but by schema density. In my environment, 56 MCP tools measured 52,942 characters and 14,011 tokens on o200k_base. The 30,000 to 50,000 token assumption in this article (for 50 tools) therefore varies widely by sample, and the decision to adopt routing should be made on the measured token count of your own tool set.
1. Measurement conditions and values
/home/jw/.mcp-proxy-cache/schemas.json,/home/jw/.hermes/cache/mcp_schema_cache.json, recorded 2026-09-29 to 30)ensure_ascii=FalseJSON serializationDistribution: average 250.2 tokens per tool (945 characters), median 209.5, minimum 92, maximum 730. Accumulated from the largest down: the top 5 take 3,166 tokens, 20 take 7,766 tokens, 50 take 13,368 tokens. A flat average puts 50 tools at 12,510 tokens.
2. Why my numbers differ from the article's
The cost of one tool is governed almost entirely by its name and description length. In this sample the maximum of 730 tokens (
connector_mcp_servers) and the minimum of 92 differ by 8x. At an average of 250 tokens, reaching 30,000 tokens would require about 120 tools. The 30,000 to 50,000 figure in the article is therefore less a general value for "50 tools" and more likely a sample weighted toward verbose schemas, or a total that also adds the system prompt and skill index. Stating which one it is would improve reproducibility.3. This is what overlapping tools actually look like
Two browser automation stacks are cached at the same time in this environment. Of the 56 tools, 32 (about 57 percent) are browser-related.
The same pattern as the "four search tools" case in the article is reproduced here in the browser stack. When only the prefix (
browser_*) differs and the descriptions are similar, the model still keeps both as candidates, so deduplicating description embeddings matters more than naming rules. If the function is the same, exposing a single server reduces both cost and confusion.4. Metrics to measure alongside routing
5. Limitations
The sample is 56 tools across 9 MCP servers on a single host, encoded with o200k_base. Per-tool cost changes with the wording and language of the descriptions and with the server implementation. The A/B/C grades in the article already carry the caveat of a single experiment in the author's environment, so the two results are better read as samples of different density than as a contradiction.