--- title: "The 21 Latest AI Agent Skills, Part 7: Running and Deploying Local Models Yourself" date: 2026-10-01 model: hermes-agent category: guide summary: The final part of the 21 agent skills. It covers GGUF quantization selection, local inference with llama.cpp, high-throughput serving with vLLM, and the Hugging Face CLI. tags: agent-skills, llama-cpp, vllm, huggingface, quantization, guide author_type: human --- Through Part 6 we handled documents. The final Part 7 is the model itself. Run it directly, deploy it, and fetch it. Each of the three skills handles one of those stages. From here on, the character differs from the other parts. The earlier skills were procedures made of text. This part involves hardware. How many gigabytes of VRAM are left changes the answer. ## 1. llama-cpp (v2.1.2) The `mlops/inference` category. Built by Orchestra Research. Its version, 2.1.2, is the highest. ```text llama.cpp local GGUF inference + HF Hub model discovery. ``` ### When to use it ```text - Run local models on CPU, Apple Silicon, CUDA, ROCm, Intel GPU - Find the right GGUF in a specific Hugging Face repository - Build llama-server or llama-cli commands from the Hub - Search Hub models that already support llama.cpp - Enumerate the .gguf files and sizes in a repository - Choose a Q4/Q5/Q6/IQ variant to fit RAM or VRAM ``` The fifth item is why this skill was made. Choosing a model is not the hard part — choosing the quantization level is. ### Model discovery starts from the URL The skill gives an instruction. ```text Before using hf, Python, or custom scripts, prefer the URL path. ``` Step 1 is search. ```text https://huggingface.co/models?apps=llama.cpp&sort=trending ``` To find a model family, append `search=`. If there is a size limit, use something like `num_parameters=min:0,max:24B`. Step 2 is the repository's local app view. ```text https://huggingface.co/?local-app=llama.cpp ``` Step 3 onward is the essence of this skill. ```text If a local app snippet appears, treat it as truth. - Copy the exact llama-server or llama-cli command as-is - Report the recommended quantization exactly as shown by HF ``` Here "copy as-is" is important. The moment a person reads it and rewrites it their own way, it is wrong. The command HF computed for that repository is right there. Step 4 reads the hardware compatibility section. ```text In the Hardware compatibility section, prefer the exact quantization labels and sizes. Prefer the repo-specific labels over generic tables (things like UD-Q4_K_M, IQ4_NL_XL). If not visible, say so and move on to the tree API and general guides. ``` Step 5 confirms actual existence with the tree API. ```text https://huggingface.co/api/models//tree/main?recursive=true Keep only entries where type is "file" and path ends in .gguf. Use path and size as the truth for filename and byte size. Separate quantization checkpoints, mmproj-*.gguf projector files, and BF16/shard files. ``` Here separating `mmproj` is the key. If you look at a vision model repository, several projector files show up. These are not the main model files. If you download the projector too, the size doubles and the agent does not know it. The skill explicitly says to separate them. Step 6 reconstructs when no snippet is visible. ```bash # Short form llama-server -hf : # With the exact file llama-server --hf-repo --hf-file ``` Step 7 is the final guard. ```text Only suggest converting Transformers weights when the repository does not already expose GGUF. ``` Conversion is slow and quality drops depending on the environment. There is no reason to convert a repository that already has GGUF. ### The order of choosing a quantization The skill specifies the priority. The Hub page first, the general heuristic second. ```text - Prefer the exact quantization HF marks as compatible with the user's hardware profile - For general chat, start at Q4_K_M - For code or technical work, prefer Q5_K_M or Q6_K if memory allows - If RAM is very tight, Q3_K_M, IQ variants, Q2 variants. Only when the user explicitly prioritizes fit over quality - For multimodal repositories, mention mmproj-*.gguf separately. The projector is not the main model file - Do not normalize repo-specific labels. If the page says UD-Q4_K_M, report UD-Q4_K_M ``` The last line saves time in practice. If you write `UD-Q4_K_M` as just `Q4_K_M`, you end up looking for a file that does not exist. It is a variant with raised precision via an underscore, so performance is better, but normalizing the name loses that benefit. One thing to point out here. v2.1.2 of this skill has an elaborate discovery procedure. When the agent is told "download a 9B model with Q4_K_M," it easily grabs just anything. This skill is the device that stops that. ## 2. serving-llms-vllm (v1.0.1) The `mlops/inference` category. Orchestra Research. ```text vLLM: high-throughput LLM serving, OpenAI API, quantization. ``` ### When to use it ```text When deploying a production LLM API, when optimizing inference latency and throughput, when serving a model with limited GPU memory ``` ### Why the throughput differs ```text Thanks to PagedAttention (block-based KV cache) and continuous batching (a technique that mixes prefill and decode requests), vLLM achieves 24x higher throughput than standard transformers. ``` The two techniques each solve a different bottleneck. PagedAttention manages the KV cache not as contiguous memory but as blocks, eliminating fragmentation and raising cache utilization. Continuous batching ignores the idea that the next request starts only when one request finishes, and inserts a new request next to a request in decode. The more traffic comes in, the more the gap widens. ### Basic usage ```bash pip install vllm ``` Offline inference. ```python from vllm import LLM, SamplingParams llm = LLM(model="meta-llama/Meta-Llama-3-8B-Instruct") sampling = SamplingParams(temperature=0.7, max_tokens=256) outputs = llm.generate(["Explain quantum computing"], sampling) print(outputs[0].outputs[0].text) ``` OpenAI-compatible server. ```bash vllm serve meta-llama/Meta-Llama-3-8B-Instruct ``` ```python from openai import OpenAI client = OpenAI(base_url='http://localhost:8000/v1', api_key='EMPTY') print(client.chat.completions.create( model='meta-llama/Meta-Llama-3-8B-Instruct', messages=[{'role': 'user', 'content': 'Hello!'}] ).choices[0].message.content) ``` `api_key='EMPTY'` stands out. It accepts any string without authentication. This is for local and internal networks only. The skill says so and puts the production deployment example first. ### Which one to use when This skill lays it out in a table. This table is the most practical thing in this part. | Situation | Tool | | --- | --- | | Production API deployment (100+ requests/sec) | vLLM | | OpenAI-compatible endpoint | vLLM | | Need a large model but GPU memory is short | vLLM | | Multi-user application | vLLM | | CPU/edge inference, single user | llama.cpp | | Research, prototyping, one-off generation | transformers | | NVIDIA-only, absolute maximum performance | TensorRT-LLM | | Already using the Hub ecosystem | TGI | The skills covered in the previous six parts are in the table. llama.cpp is on the CPU/edge side, and transformers is on the research side. So choosing this one skill determines where the other two sit. ### Out-of-memory is the most common problem ```bash vllm serve MODEL \ --gpu-memory-utilization 0.7 \ --max-model-len 4096 ``` This is the solution when VRAM is insufficient and the model will not load. Lower the GPU memory utilization cap and reduce the context length. Context length directly determines the KV cache size, so lowering the value here has the biggest effect. You can also use quantization (GPTQ, AWQ, FP8). The skill includes it in the supported list. ### Hardware requirements ```text - Small (7B–13B): 1x A10 (24GB) or 1x A100 (40GB) - Medium (30B–40B): 2x A100 (40GB), using tensor parallelism - Large (70B+): 4x A100 (40GB) or 2x (80GB), using AWQ/GPTQ ``` Supporting platforms are NVIDIA first, followed by AMD ROCm, Intel GPU, and TPU. In other words, if it is not NVIDIA, the case for vLLM is weak. In that case, llama.cpp is the right choice. ## 3. huggingface-hub (v1.0.1) The `mlops` category. Built by Hugging Face itself. It is the only one of the three that belongs to official documentation. ```text HuggingFace hf CLI: search/download/upload models, datasets. ``` The install is short too. ```bash curl -LsSf https://hf.co/cli/install.sh | bash -s ``` And the skill warns. ```text IMPORTANT: the hf command now replaces the deprecated huggingface-cli. ``` If you use old scripts as-is, the command is not found. On top of that, many environments still have `huggingface-cli download` installed. This distinction is confusing, and it explicitly says to standardize on the new command. For auth it recommends the `HF_TOKEN` environment variable or the `--token` flag. ### Core commands ```text hf download REPO_ID download files from the Hub hf upload REPO_ID upload files/folders (single commit recommended, large directories use resumable upload) hf upload-large-folder [Deprecated] — use hf upload hf sync sync between a local directory and a bucket hf env / hf version check environment/version ``` It is important that `upload-large-folder` is marked deprecated. If you follow an old guide, you will use that command. ### Repository management ```text create / delete create/permanently delete a repository duplicate duplicate a model/dataset/Space under a new ID move move between namespaces branch / tag Git-like reference management delete-files delete specific files by pattern ``` `duplicate` is especially useful. It is used when creating a derivative model based on someone else's repository. It is clearer than forking. ### Specialized hub features ```text Datasets: hf datasets list, info, parquet SQL: hf datasets sql SQL — run raw SQL with DuckDB against a dataset parquet URL Models: hf models list, info Papers: hf papers ls — today's papers ``` `hf datasets sql` stands out. It throws SQL directly at a parquet URL. You can query a dataset without downloading it locally. This is fast when scanning a large dataset. Discussions and PRs. ```text list, create, info, comment, close, reopen, rename diff view PR changes merge finalize a PR ``` It handles through the CLI the flow where people leave issues on a model repository you created. Infrastructure and compute. ```text Endpoints: deploy, pause, resume, scale-to-zero, catalog Jobs: hf jobs uv (run Python with inline dependencies), stats (resource monitoring) Spaces: dev-mode, hot reload ``` `scale-to-zero` is practical. It is the method of paying only while you use an Inference Endpoint. It stays off the rest of the time. Storage and automation. ```text Buckets: create, cp, mv, rm, sync Cache: list, prune (remove detached revisions), verify (checksum) Webhooks: create, watch, enable/disable Collections: add-item, update, list ``` `cache prune` is useful for disk cleanup. You do it later, but once you download several models the cache becomes tens of GB. ### Global flags ```text --format json machine-readable output -q / --quiet output only the ID ``` For an agent, `--format json` is almost mandatory. If you try to parse a human-readable log, the format breaks. And reducing output volume saves context. There are extensions too. ```text Extensions: hf extensions install REPO_ID (extend CLI functionality from a GitHub repository) ``` ## Part 7 Summary | Skill | Version | What it does | | --- | --- | --- | | llama-cpp | 2.1.2 | GGUF quantization selection and local llama.cpp inference, URL-first discovery | | serving-llms-vllm | 1.0.1 | High-throughput serving with PagedAttention and continuous batching | | huggingface-hub | 1.0.1 | Search, download, upload models and datasets with the hf CLI | The three form a single line from local to deployment. Find it on the Hub (3), choose the quantization level and download (1), and if needed deploy with high throughput (2). And there is one way this part differs from the others. The other six parts covered "how to do it well," and this part covered "what to choose." Choose the quantization level, choose whether it is vLLM or llama.cpp, and check whether the hardware supports it. There is no single right answer in technical choices. There is only the cost of a wrong choice. ## The Full List of 21 We have now covered all 21. Looking back at the series and regrouping, it is this. ```text Part 1 Skill basics hermes-agent-skill-authoring, plan, computer-use Part 2 Coding delegation delegate-coding, delegate-debugging, test-driven-development Part 3 Debugging systematic-debugging, python-debugpy, node-inspect-debugger Part 4 Quality management requesting-code-review, codebase-inspection, code-reference Part 5 Research deep-web-investigation, grounded-citations, research-paper-writing Part 6 Documents/notes notion, obsidian, nano-pdf Part 7 Model operations llama-cpp, serving-llms-vllm, huggingface-hub ``` Part 1 pinned down the definition of what a skill is, and Part 7 came all the way to how to choose a model. The five parts in between are all the hands-on work in between. Delegating, verifying, fixing, checking, citing, and handling documents. One thing common to all of them shows here. These skills are all built to preempt the points where an agent causes an accident. The instructions inside the diff in Part 4, the uncited claim with no ledger in Part 5, the unshared page that spits out a 404 in Part 6. All of them are "points the agent passes through without knowing." Blocking those points one by one is what a skill is.