The 21 Latest AI Agent Skills, Part 7: Running and Deploying Local Models Yourself
Through Part 6 we handled documents. The final Part 7 is the model itself. Run it directly, deploy it, and fetch it. Each of the three skills handles one of those stages.
From here on, the character differs from the other parts. The earlier skills were procedures made of text. This part involves hardware. How many gigabytes of VRAM are left changes the answer.
1. llama-cpp (v2.1.2)
The mlops/inference category. Built by Orchestra Research. Its version, 2.1.2, is the highest.
llama.cpp local GGUF inference + HF Hub model discovery.
When to use it
- Run local models on CPU, Apple Silicon, CUDA, ROCm, Intel GPU
- Find the right GGUF in a specific Hugging Face repository
- Build llama-server or llama-cli commands from the Hub
- Search Hub models that already support llama.cpp
- Enumerate the .gguf files and sizes in a repository
- Choose a Q4/Q5/Q6/IQ variant to fit RAM or VRAM
The fifth item is why this skill was made. Choosing a model is not the hard part โ choosing the quantization level is.
Model discovery starts from the URL
The skill gives an instruction.
Before using hf, Python, or custom scripts, prefer the URL path.
Step 1 is search.
https://huggingface.co/models?apps=llama.cpp&sort=trending
To find a model family, append search=<term>. If there is a size limit, use something like num_parameters=min:0,max:24B.
Step 2 is the repository's local app view.
https://huggingface.co/<repo>?local-app=llama.cpp
Step 3 onward is the essence of this skill.
If a local app snippet appears, treat it as truth.
- Copy the exact llama-server or llama-cli command as-is
- Report the recommended quantization exactly as shown by HF
Here "copy as-is" is important. The moment a person reads it and rewrites it their own way, it is wrong. The command HF computed for that repository is right there.
Step 4 reads the hardware compatibility section.
In the Hardware compatibility section, prefer the exact quantization labels and sizes.
Prefer the repo-specific labels over generic tables (things like UD-Q4_K_M, IQ4_NL_XL).
If not visible, say so and move on to the tree API and general guides.
Step 5 confirms actual existence with the tree API.
https://huggingface.co/api/models/<repo>/tree/main?recursive=true
Keep only entries where type is "file" and path ends in .gguf.
Use path and size as the truth for filename and byte size.
Separate quantization checkpoints, mmproj-*.gguf projector files, and BF16/shard files.
Here separating mmproj is the key. If you look at a vision model repository, several projector files show up. These are not the main model files. If you download the projector too, the size doubles and the agent does not know it. The skill explicitly says to separate them.
Step 6 reconstructs when no snippet is visible.
# Short form
llama-server -hf <repo>:<QUANT>
# With the exact file
llama-server --hf-repo <repo> --hf-file <filename.gguf>
Step 7 is the final guard.
Only suggest converting Transformers weights when the repository does not already expose GGUF.
Conversion is slow and quality drops depending on the environment. There is no reason to convert a repository that already has GGUF.
The order of choosing a quantization
The skill specifies the priority. The Hub page first, the general heuristic second.
- Prefer the exact quantization HF marks as compatible with the user's hardware profile
- For general chat, start at Q4_K_M
- For code or technical work, prefer Q5_K_M or Q6_K if memory allows
- If RAM is very tight, Q3_K_M, IQ variants, Q2 variants. Only when the user explicitly prioritizes fit over quality
- For multimodal repositories, mention mmproj-*.gguf separately. The projector is not the main model file
- Do not normalize repo-specific labels. If the page says UD-Q4_K_M, report UD-Q4_K_M
The last line saves time in practice. If you write UD-Q4_K_M as just Q4_K_M, you end up looking for a file that does not exist. It is a variant with raised precision via an underscore, so performance is better, but normalizing the name loses that benefit.
One thing to point out here. v2.1.2 of this skill has an elaborate discovery procedure. When the agent is told "download a 9B model with Q4_K_M," it easily grabs just anything. This skill is the device that stops that.
2. serving-llms-vllm (v1.0.1)
The mlops/inference category. Orchestra Research.
vLLM: high-throughput LLM serving, OpenAI API, quantization.
When to use it
When deploying a production LLM API, when optimizing inference latency and throughput,
when serving a model with limited GPU memory
Why the throughput differs
Thanks to PagedAttention (block-based KV cache) and continuous batching
(a technique that mixes prefill and decode requests), vLLM achieves
24x higher throughput than standard transformers.
The two techniques each solve a different bottleneck. PagedAttention manages the KV cache not as contiguous memory but as blocks, eliminating fragmentation and raising cache utilization. Continuous batching ignores the idea that the next request starts only when one request finishes, and inserts a new request next to a request in decode. The more traffic comes in, the more the gap widens.
Basic usage
pip install vllm
Offline inference.
from vllm import LLM, SamplingParams
llm = LLM(model="meta-llama/Meta-Llama-3-8B-Instruct")
sampling = SamplingParams(temperature=0.7, max_tokens=256)
outputs = llm.generate(["Explain quantum computing"], sampling)
print(outputs[0].outputs[0].text)
OpenAI-compatible server.
vllm serve meta-llama/Meta-Llama-3-8B-Instruct
from openai import OpenAI
client = OpenAI(base_url='http://localhost:8000/v1', api_key='EMPTY')
print(client.chat.completions.create(
model='meta-llama/Meta-Llama-3-8B-Instruct',
messages=[{'role': 'user', 'content': 'Hello!'}]
).choices[0].message.content)
api_key='EMPTY' stands out. It accepts any string without authentication. This is for local and internal networks only. The skill says so and puts the production deployment example first.
Which one to use when
This skill lays it out in a table. This table is the most practical thing in this part.
| Situation | Tool |
|---|---|
| Production API deployment (100+ requests/sec) | vLLM |
| OpenAI-compatible endpoint | vLLM |
| Need a large model but GPU memory is short | vLLM |
| Multi-user application | vLLM |
| CPU/edge inference, single user | llama.cpp |
| Research, prototyping, one-off generation | transformers |
| NVIDIA-only, absolute maximum performance | TensorRT-LLM |
| Already using the Hub ecosystem | TGI |
The skills covered in the previous six parts are in the table. llama.cpp is on the CPU/edge side, and transformers is on the research side. So choosing this one skill determines where the other two sit.
Out-of-memory is the most common problem
vllm serve MODEL \
--gpu-memory-utilization 0.7 \
--max-model-len 4096
This is the solution when VRAM is insufficient and the model will not load. Lower the GPU memory utilization cap and reduce the context length. Context length directly determines the KV cache size, so lowering the value here has the biggest effect.
You can also use quantization (GPTQ, AWQ, FP8). The skill includes it in the supported list.
Hardware requirements
- Small (7Bโ13B): 1x A10 (24GB) or 1x A100 (40GB)
- Medium (30Bโ40B): 2x A100 (40GB), using tensor parallelism
- Large (70B+): 4x A100 (40GB) or 2x (80GB), using AWQ/GPTQ
Supporting platforms are NVIDIA first, followed by AMD ROCm, Intel GPU, and TPU. In other words, if it is not NVIDIA, the case for vLLM is weak. In that case, llama.cpp is the right choice.
3. huggingface-hub (v1.0.1)
The mlops category. Built by Hugging Face itself. It is the only one of the three that belongs to official documentation.
HuggingFace hf CLI: search/download/upload models, datasets.
The install is short too.
curl -LsSf https://hf.co/cli/install.sh | bash -s
And the skill warns.
IMPORTANT: the hf command now replaces the deprecated huggingface-cli.
If you use old scripts as-is, the command is not found. On top of that, many environments still have huggingface-cli download installed. This distinction is confusing, and it explicitly says to standardize on the new command.
For auth it recommends the HF_TOKEN environment variable or the --token flag.
Core commands
hf download REPO_ID download files from the Hub
hf upload REPO_ID upload files/folders (single commit recommended, large directories use resumable upload)
hf upload-large-folder [Deprecated] โ use hf upload
hf sync sync between a local directory and a bucket
hf env / hf version check environment/version
It is important that upload-large-folder is marked deprecated. If you follow an old guide, you will use that command.
Repository management
create / delete create/permanently delete a repository
duplicate duplicate a model/dataset/Space under a new ID
move move between namespaces
branch / tag Git-like reference management
delete-files delete specific files by pattern
duplicate is especially useful. It is used when creating a derivative model based on someone else's repository. It is clearer than forking.
Specialized hub features
Datasets: hf datasets list, info, parquet
SQL: hf datasets sql SQL โ run raw SQL with DuckDB against a dataset parquet URL
Models: hf models list, info
Papers: hf papers ls โ today's papers
hf datasets sql stands out. It throws SQL directly at a parquet URL. You can query a dataset without downloading it locally. This is fast when scanning a large dataset.
Discussions and PRs.
list, create, info, comment, close, reopen, rename
diff view PR changes
merge finalize a PR
It handles through the CLI the flow where people leave issues on a model repository you created.
Infrastructure and compute.
Endpoints: deploy, pause, resume, scale-to-zero, catalog
Jobs: hf jobs uv (run Python with inline dependencies), stats (resource monitoring)
Spaces: dev-mode, hot reload
scale-to-zero is practical. It is the method of paying only while you use an Inference Endpoint. It stays off the rest of the time.
Storage and automation.
Buckets: create, cp, mv, rm, sync
Cache: list, prune (remove detached revisions), verify (checksum)
Webhooks: create, watch, enable/disable
Collections: add-item, update, list
cache prune is useful for disk cleanup. You do it later, but once you download several models the cache becomes tens of GB.
Global flags
--format json machine-readable output
-q / --quiet output only the ID
For an agent, --format json is almost mandatory. If you try to parse a human-readable log, the format breaks. And reducing output volume saves context.
There are extensions too.
Extensions: hf extensions install REPO_ID (extend CLI functionality from a GitHub repository)
Part 7 Summary
| Skill | Version | What it does |
|---|---|---|
| llama-cpp | 2.1.2 | GGUF quantization selection and local llama.cpp inference, URL-first discovery |
| serving-llms-vllm | 1.0.1 | High-throughput serving with PagedAttention and continuous batching |
| huggingface-hub | 1.0.1 | Search, download, upload models and datasets with the hf CLI |
The three form a single line from local to deployment. Find it on the Hub (3), choose the quantization level and download (1), and if needed deploy with high throughput (2).
And there is one way this part differs from the others. The other six parts covered "how to do it well," and this part covered "what to choose." Choose the quantization level, choose whether it is vLLM or llama.cpp, and check whether the hardware supports it. There is no single right answer in technical choices. There is only the cost of a wrong choice.
The Full List of 21
We have now covered all 21. Looking back at the series and regrouping, it is this.
Part 1 Skill basics hermes-agent-skill-authoring, plan, computer-use
Part 2 Coding delegation delegate-coding, delegate-debugging, test-driven-development
Part 3 Debugging systematic-debugging, python-debugpy, node-inspect-debugger
Part 4 Quality management requesting-code-review, codebase-inspection, code-reference
Part 5 Research deep-web-investigation, grounded-citations, research-paper-writing
Part 6 Documents/notes notion, obsidian, nano-pdf
Part 7 Model operations llama-cpp, serving-llms-vllm, huggingface-hub
Part 1 pinned down the definition of what a skill is, and Part 7 came all the way to how to choose a model. The five parts in between are all the hands-on work in between. Delegating, verifying, fixing, checking, citing, and handling documents.
One thing common to all of them shows here. These skills are all built to preempt the points where an agent causes an accident. The instructions inside the diff in Part 4, the uncited claim with no ledger in Part 5, the unshared page that spits out a 404 in Part 6. All of them are "points the agent passes through without knowing." Blocking those points one by one is what a skill is.
AI Knowledge Hub