The Three Eras of Agent Extensibility, from ChatGPT Plugins to MCP to Agent Skills
Every time you try to make an agent do something new, you hit the same wall. The model only knows what it already knows. You want to attach a tool that exists only inside your company, or teach it how your team deploys, or make it write code in your own style. That is why extension points appeared around agents. Those points have been swapped three times.
1. The ChatGPT Plugin era β extensibility meant code
ChatGPT plugins in early 2023 were exactly honest. A plugin was a program. You stood up a server on the outside, attached authentication, declared an API spec, and ChatGPT called that server and spliced the plugin response into its answer.
What got extended was capability. "This server tells the weather." "This server books accommodation." The server did what the model could not. Web search, code execution, image generation β all followed this structure.
Looking back, the design of that era feels exhausted but consistent. There was a plugin manifest β a JSON schema β and function signatures were declared inside it. There were types. That design is actually still valid today.
But there was a clear wall. To extend, you had to program. To attach a weather tool, you had to stand up a server, issue an HTTPS certificate, and build an OAuth consent screen. A person who could not program needed two days to add one agent tool. And the most absurd part: why you built that server was already written in the code. The difference between writing down that your company has a three-step approval process and turning it into an HTTP endpoint was not large.
2. The MCP era β the extension point was still code, but installation got easier
In late 2024, Anthropic released the Model Context Protocol and the direction bent. The direction only. The extension point was still code. An MCP server is still a program, and it exposes tools, resources, and prompts as they are.
What changed was the connection method. Instead of building an auth screen and spec for every plugin, one standard protocol finished the job. You put one server address into an MCP client and the tools appeared. It swapped a "plugin architecture" for a "standard socket." This was a genuine improvement, and a standard effectively settled in place.
Something interesting happened here. Because MCP made extension easier, people began questioning what needed extending. As models got good enough, requests like this started arriving:
"Explain why yesterday's test failure happened"
The agent could not swap servers. It read and still produced no answer. Was the response format weird? No. Reading a failed test, interpreting the stack trace correctly, and saying which part of the code structure the failure came from β that requires knowledge. The model was not unintelligent because it did not know the pattern. Nobody had written down what happened in that project.
3. The Agent Skills era β extensibility became knowledge
From the second half of 2025, Claude Skills arrived and the direction bent again, for real.
ChatGPT Plugin: extension adds capability
MCP: connects capabilities via a standard
Agent Skills: extension adds knowledge
A skill is neither a server nor code. It is one markdown file. I scanned all 98 skills in Hermes Agent and the result was clear. Every one was pure markdown with YAML frontmatter, and the executable files were only Python or shell. More importantly, there was no code inside that calls a model. A field that specifies a model does not exist.
The frontmatter looks like this:
---
name: data-preprocessing
description: Use when reading CSV or Excel files and cleaning missing values
version: 1.0.0
platforms: [linux, macos, windows]
license: MIT
prerequisites: [pandas, openpyxl]
---
Of the 98, 85 had metadata, 84 had license, and 88 had platforms. Why does this matter? Because this is a completely portable unit. You copy a skill file into another agent and it works as-is. When I checked portability across 21 selected skills, only 2 were Hermes-specific. The other 19 work regardless of the agent name.
The description is the real interface
The most interesting part of the skill mechanism is this. The description field is effectively the call trigger. If it says "use whenβ¦", the agent finds and invokes it. If it is vague, the skill is never invoked even though it exists.
In other words, the interface shifted from a JSON schema to natural-language matching. A plugin manifest forced name, parameters, and type. A skill explains intent in a sentence. It works because the model interprets natural language. In 2023, who could have imagined a plugin that relied on description: "use whenβ¦".
What makes skills powerful
The most fun part of reading skills was how often the way they work is an instruction. One line from the TDD skill:
If you didn't watch the test fail, you don't know if it tests the right thing.
This is not code. It is an instruction to write the test first and watch it actually fail. Yet that instruction works more accurately than a program, because an interpreter cannot enforce "do not write code until you have seen the test fail."
Another one. The code-reference skill is this:
When referencing code, libraries, or APIs, do not guess. Search official documentation first.
Not code, not a script. A 15-character principle: "when you do not know, do not guess β find the official docs." Seeing why this became a skill makes the era shift sharp. The skill was needed because the model has search capability and the problem is that it does not search. What solved the problem was not code, it was instruction.
4. What differed across the three eras
| Plugin (2023) | MCP (2024) | Skills (2025) | |
|---|---|---|---|
| What is extended | capability | capability (standardized) | knowledge |
| Implementation language | server code | server code | markdown |
| Interface | JSON manifest | JSON schema | natural-language description |
| Authentication | OAuth built manually | OAuth built manually | none |
| Installation | server deployment | server deployment | file copy |
| Reusability | low | medium (protocol standard) | high (model-agnostic) |
| What you need | deploy + auth + server | server | a text file |
The column to watch is authentication. Plugins and MCP constantly collide with auth. OAuth flows, token management, scope design. Skills, on the other hand, need nothing. You put the file in a folder and that is it. This difference is what actually split adoption. I think it is the real driver of each era shift. Not technical superiority, but ease. MCP did not fully replace plugins for the same reason. The standard was right, but you still had to stand up a server.
5. Problems that remain
Ending on "skills are the best" would be less interesting, and honestly the downsides are clear.
There are no types. MCP tools have parameter types, so wrong calls surface as errors. Skills are natural language, so misinterpretation happens and sometimes no skill gets invoked at all.
Overlap appears. With 98 skills, similar descriptions come in pairs. One internal skill has to be chosen, and the choice is the model's judgment, not a human's.
Execution is deferred. Skills change how you think but cannot execute immediately. Anything that needs to run code still goes through MCP or scripts. So Hermes hosts skills and MCP together. The roles differ. Skills tell you "how to do it," MCP provides "what you can do."
Closing
Across three era shifts, one pattern holds. The stronger the model gets, the less code extensibility requires. In 2023, adding weather data required standing up a server. In 2024, teaching a process rule required building an MCP server. In 2025, both are one markdown file.
One question remains. How much knowledge cannot be expressed as code? TDD is a principle that cannot be reduced to true/false, so markdown suited it. Process rules too. Then could something appear that text cannot carry? Skills went as far as text can go, and what comes next might be another medium.
It does not appear to have arrived yet.
AI Knowledge Hub