--- title: 'Six Thousand Five Hundred Tokens for One "Hello": The Real Face of AI Agent Feature Bloat, Examined by Actual Hermes and Claude Skill Names' date: 2026-09-29 model: opencode category: knowhow summary: "Weighing Hermes's 40-plus built-in skills against Claude's pptx/xlsx/docx/pdf skills shows that the battleground is not 'we support 40 skills' but 'the real context cost of a single hello.'" tags: AIagents,skills,tokens,Claude,Hermes,featurebloat --- Whenever you read the "major update" news dropping out of the AI industry every week, the first reaction is a sigh. You have not even finished using the features added last week, and today another release boasts dozens of new skills and autonomous agents wrapped in grandiose fanfare. The honest sentiment of someone who uses AI every day comes down to a single sentence. **"Why are you cramming in so many features? I only ever use a few, and all that happens in the background is my tokens burning."** Let us trace the reality of the "agent feature bloat" that Big Tech and startups are competing over, and the "hidden token tax" that keeps getting billed where the user cannot see it, using actual agent and skill names. ## 1. Feature bloat is the "dead feature" of the AI era According to Standish Group, a long-standing research institution in software engineering, **45% of all features in traditional commercial software are never used at all, and 19% are used once in a while**. In other words, 64% of features are effectively bloatware. In the AI agent ecosystem this phenomenon is more serious. A skill is not something that is merely "nice to have." It is **a contract permanently injected into the prompt as a schema**. | Metric | Figure and reality | Note | | --- | --- | --- | | Share of dead features | **64%** of all shipped features (45% never used + 19% barely used) | Standish Group CHAOS Report | | Core use cases of general users | More than 80% concentrated in the **top 3 to 4**: simple questions, summarization, translation, sentence correction | analysis of large-scale conversation logs | | Adoption rate of complex multi-step skills | Advanced skills requiring multiple tool calls are used **less than 5%** of the time | B2B/B2C SaaS usability analysis | What users need is the performance of "understanding my intent in one go and answering it." It is not a Swiss army knife holding a hundred tools. ## 2. Named comparison 1: Hermes — what "40-plus built-in skills" actually is Hermes is an open-source agent released by Nous Research. It advertises more than 40 built-in skills. What are those skills actually named and shaped like? **Skills registered as slash commands (SKILL.md based)** - `/gif-search` — GIF search - `/axolotl` — axolotl server operation and management - `/github-pr-workflow` — the PR authoring flow - `/songsee` — music-related visualization and processing - `/excalidraw` — diagram generation - `/plan` — planning - `/ocr-and-documents` — OCR and document processing **Multi-skill stacking** — Hermes does not force you to pick just one. You can chain them. ``` /github-pr-workflow /test-driven-development fix issue #123 and open a PR ``` **Bundles** — related skills grouped into predefined combinations. For example `backend-dev` is the combination of `github-code-review` + `test-driven-development` + `github-pr-workflow`. Type it once and the instructions of all three skills come in at once. **Fallback skills** — alternative paths are defined for environments that lack a toolset. For instance `duckduckgo-search` is declared with `fallback_for_toolsets: [web]`, so if there is no web tool it substitutes search instead. These defensive skills have a low real usage rate, yet they still occupy schema space even for a single "hello." **Extension installs** — external skills can be added with something like `hermes skills install official/security/1password`. In other words, 40-plus is only the starting point. The schema keeps growing. To its credit, Hermes is honest in design here. It uses a **progressive disclosure** structure: it first puts only the list of skills (a summary) into context, and the actual body is read when that skill is triggered. The load order is list, then body, then reference files. But however tidy the structure, the volume of the tool schemas themselves does not disappear. ## 3. Named comparison 2: Claude — what the `pptx`/`xlsx`/`docx`/`pdf` skills give you On the Claude side the names are equally clear. - **Built-in skill IDs**: `pptx`, `xlsx`, `docx`, `pdf` — preconfigured skills offered through the API, claude.ai, AWS Bedrock, and Microsoft Foundry. - **On the Claude Code side**: these built-in document skills are not shipped as-is, and you have to assemble custom skills yourself. - **Open source repository**: `anthropics/skills` publishes the `docx`, `pdf`, `pptx`, and `xlsx` skills. The installation path is fixed as well. ``` /plugin marketplace add anthropics/skills /plugin install document-skills@anthropic-agent-skills /example-skills@anthropic-agent-skills ``` Claude's skill system saves tokens with **three levels of progressive disclosure**. | Level | Content | Token size | When it loads | | --- | --- | --- | --- | | Level 1 | Skill name and description metadata | about **100 tokens** per skill | **Always**, for the whole session | | Level 2 | The `SKILL.md` body | **under 5 thousand tokens** | Once, at the moment of trigger | | Level 3 | Reference files and scripts | variable | On demand. Scripts do not put code in context, only their **output** is brought back | In other words, the structure is "Level 1 is a permanent cost proportional to the number of skills, Level 2 is a cost only when used." Add the `paths` frontmatter of `.claude/rules/` and even the rules activate at the moment the file is actually Read, so on unrelated work the instructions do not occupy context. ## 4. So why does "hello" eat 6,500 tokens Adding up the numbers as they are. - Base system prompt: **1,000 to 2,000 tokens** - Skill schemas (JSON definitions): once you pass 10 skills, **2,000 to 5,000 tokens** - Instruction files and rules (`.claude/rules/`, agent instructions): when nested, **another 2,000 to 5,000 tokens** ``` [What the user actually typed] "How is the weather in Gwangju today?" (about 10 tokens) [The context the model actually received] system prompt + many skill schemas + exception-handling guide (about 6,510 tokens) ``` With a 40-plus skill structure like Hermes, Level 1 metadata alone is 40 × 100 = **4,000 tokens** permanently in context. Add the system prompt and the instruction files and the user has thrown in 10 tokens while 6,000 to 7,000 tokens have already been billed. So the real measurement is not the number of skills but **"the actual context cost of a single hello."** ## 5. The technical diet: five calculations Reducing feature bloat is not a marketing problem but an engineering one. The following five things can simply be calculated. 1. **Constant tokens**: number of skills × Level 1 metadata + system prompt + always-loaded rules. This value is the base tax of every request. 2. **Trigger cost**: the Level 2 body size of the 3 to 4 skills you actually use often. Narrowing the frequently used paths with `paths` makes the rest zero. 3. **Idle cost**: constant tokens keep accumulating even in a session where nothing is happening. This is the real cost of an agent left with its window open. 4. **Conflict cost**: the more skills and rules there are, the more conflicting instructions appear, and the model forgets important instructions or mixes them together. Hallucination and instruction non-compliance come from here. 5. **Reuse rate**: the quality of the result is judged only by the final artifact. If more orchestration layers make the output worse, all of that cost was pure waste. ## Conclusion: the battle is over the cost of "one hello," not 500 skills The battleground of AI agents from here is not "we support 500 skills." Whether Hermes pushes 40 of them or Claude offers four document skills, as long as the user does not select that schema, the schema is a pure tax. An agent where one "hello" costs 6,500 tokens of real context, versus an agent where it costs 650 tokens. That sevenfold gap is the difference in real product capability. The AI agent market is coming to resemble the early smartphone apps that ate memory with all kinds of crude features. Instead of cramming in a hundred, the winner will be the one that makes the three skills users touch every day exactly 100% reliable.