The 21 Latest AI Agent Skills, Part 7: Running and Deploying Local Models Yourself

The final part of the 21 agent skills. It covers GGUF quantization selection, local inference with llama.cpp, high-throughput serving with vLLM, and the Hugging Face CLI.
Markdown sourceยทAnything to add or correct?

Through Part 6 we handled documents. The final Part 7 is the model itself. Run it directly, deploy it, and fetch it. Each of the three skills handles one of those stages.

From here on, the character differs from the other parts. The earlier skills were procedures made of text. This part involves hardware. How many gigabytes of VRAM are left changes the answer.

1. llama-cpp (v2.1.2)

The mlops/inference category. Built by Orchestra Research. Its version, 2.1.2, is the highest.


llama.cpp local GGUF inference + HF Hub model discovery.

When to use it


- Run local models on CPU, Apple Silicon, CUDA, ROCm, Intel GPU
- Find the right GGUF in a specific Hugging Face repository
- Build llama-server or llama-cli commands from the Hub
- Search Hub models that already support llama.cpp
- Enumerate the .gguf files and sizes in a repository
- Choose a Q4/Q5/Q6/IQ variant to fit RAM or VRAM

The fifth item is why this skill was made. Choosing a model is not the hard part โ€” choosing the quantization level is.

Model discovery starts from the URL

The skill gives an instruction.


Before using hf, Python, or custom scripts, prefer the URL path.

Step 1 is search.


https://huggingface.co/models?apps=llama.cpp&sort=trending

To find a model family, append search=<term>. If there is a size limit, use something like num_parameters=min:0,max:24B.

Step 2 is the repository's local app view.


https://huggingface.co/<repo>?local-app=llama.cpp

Step 3 onward is the essence of this skill.


If a local app snippet appears, treat it as truth.
- Copy the exact llama-server or llama-cli command as-is
- Report the recommended quantization exactly as shown by HF

Here "copy as-is" is important. The moment a person reads it and rewrites it their own way, it is wrong. The command HF computed for that repository is right there.

Step 4 reads the hardware compatibility section.


In the Hardware compatibility section, prefer the exact quantization labels and sizes.
Prefer the repo-specific labels over generic tables (things like UD-Q4_K_M, IQ4_NL_XL).
If not visible, say so and move on to the tree API and general guides.

Step 5 confirms actual existence with the tree API.


https://huggingface.co/api/models/<repo>/tree/main?recursive=true
Keep only entries where type is "file" and path ends in .gguf.
Use path and size as the truth for filename and byte size.
Separate quantization checkpoints, mmproj-*.gguf projector files, and BF16/shard files.

Here separating mmproj is the key. If you look at a vision model repository, several projector files show up. These are not the main model files. If you download the projector too, the size doubles and the agent does not know it. The skill explicitly says to separate them.

Step 6 reconstructs when no snippet is visible.


# Short form
llama-server -hf <repo>:<QUANT>

# With the exact file
llama-server --hf-repo <repo> --hf-file <filename.gguf>

Step 7 is the final guard.


Only suggest converting Transformers weights when the repository does not already expose GGUF.

Conversion is slow and quality drops depending on the environment. There is no reason to convert a repository that already has GGUF.

The order of choosing a quantization

The skill specifies the priority. The Hub page first, the general heuristic second.


- Prefer the exact quantization HF marks as compatible with the user's hardware profile
- For general chat, start at Q4_K_M
- For code or technical work, prefer Q5_K_M or Q6_K if memory allows
- If RAM is very tight, Q3_K_M, IQ variants, Q2 variants. Only when the user explicitly prioritizes fit over quality
- For multimodal repositories, mention mmproj-*.gguf separately. The projector is not the main model file
- Do not normalize repo-specific labels. If the page says UD-Q4_K_M, report UD-Q4_K_M

The last line saves time in practice. If you write UD-Q4_K_M as just Q4_K_M, you end up looking for a file that does not exist. It is a variant with raised precision via an underscore, so performance is better, but normalizing the name loses that benefit.

One thing to point out here. v2.1.2 of this skill has an elaborate discovery procedure. When the agent is told "download a 9B model with Q4_K_M," it easily grabs just anything. This skill is the device that stops that.

2. serving-llms-vllm (v1.0.1)

The mlops/inference category. Orchestra Research.


vLLM: high-throughput LLM serving, OpenAI API, quantization.

When to use it


When deploying a production LLM API, when optimizing inference latency and throughput,
when serving a model with limited GPU memory

Why the throughput differs


Thanks to PagedAttention (block-based KV cache) and continuous batching
(a technique that mixes prefill and decode requests), vLLM achieves
24x higher throughput than standard transformers.

The two techniques each solve a different bottleneck. PagedAttention manages the KV cache not as contiguous memory but as blocks, eliminating fragmentation and raising cache utilization. Continuous batching ignores the idea that the next request starts only when one request finishes, and inserts a new request next to a request in decode. The more traffic comes in, the more the gap widens.

Basic usage


pip install vllm

Offline inference.


from vllm import LLM, SamplingParams

llm = LLM(model="meta-llama/Meta-Llama-3-8B-Instruct")
sampling = SamplingParams(temperature=0.7, max_tokens=256)

outputs = llm.generate(["Explain quantum computing"], sampling)
print(outputs[0].outputs[0].text)

OpenAI-compatible server.


vllm serve meta-llama/Meta-Llama-3-8B-Instruct

from openai import OpenAI
client = OpenAI(base_url='http://localhost:8000/v1', api_key='EMPTY')
print(client.chat.completions.create(
    model='meta-llama/Meta-Llama-3-8B-Instruct',
    messages=[{'role': 'user', 'content': 'Hello!'}]
).choices[0].message.content)

api_key='EMPTY' stands out. It accepts any string without authentication. This is for local and internal networks only. The skill says so and puts the production deployment example first.

Which one to use when

This skill lays it out in a table. This table is the most practical thing in this part.

SituationTool
Production API deployment (100+ requests/sec)vLLM
OpenAI-compatible endpointvLLM
Need a large model but GPU memory is shortvLLM
Multi-user applicationvLLM
CPU/edge inference, single userllama.cpp
Research, prototyping, one-off generationtransformers
NVIDIA-only, absolute maximum performanceTensorRT-LLM
Already using the Hub ecosystemTGI

The skills covered in the previous six parts are in the table. llama.cpp is on the CPU/edge side, and transformers is on the research side. So choosing this one skill determines where the other two sit.

Out-of-memory is the most common problem


vllm serve MODEL \
  --gpu-memory-utilization 0.7 \
  --max-model-len 4096

This is the solution when VRAM is insufficient and the model will not load. Lower the GPU memory utilization cap and reduce the context length. Context length directly determines the KV cache size, so lowering the value here has the biggest effect.

You can also use quantization (GPTQ, AWQ, FP8). The skill includes it in the supported list.

Hardware requirements


- Small (7Bโ€“13B): 1x A10 (24GB) or 1x A100 (40GB)
- Medium (30Bโ€“40B): 2x A100 (40GB), using tensor parallelism
- Large (70B+): 4x A100 (40GB) or 2x (80GB), using AWQ/GPTQ

Supporting platforms are NVIDIA first, followed by AMD ROCm, Intel GPU, and TPU. In other words, if it is not NVIDIA, the case for vLLM is weak. In that case, llama.cpp is the right choice.

3. huggingface-hub (v1.0.1)

The mlops category. Built by Hugging Face itself. It is the only one of the three that belongs to official documentation.


HuggingFace hf CLI: search/download/upload models, datasets.

The install is short too.


curl -LsSf https://hf.co/cli/install.sh | bash -s

And the skill warns.


IMPORTANT: the hf command now replaces the deprecated huggingface-cli.

If you use old scripts as-is, the command is not found. On top of that, many environments still have huggingface-cli download installed. This distinction is confusing, and it explicitly says to standardize on the new command.

For auth it recommends the HF_TOKEN environment variable or the --token flag.

Core commands


hf download REPO_ID      download files from the Hub
hf upload REPO_ID        upload files/folders (single commit recommended, large directories use resumable upload)
hf upload-large-folder   [Deprecated] โ€” use hf upload
hf sync                 sync between a local directory and a bucket
hf env / hf version     check environment/version

It is important that upload-large-folder is marked deprecated. If you follow an old guide, you will use that command.

Repository management


create / delete      create/permanently delete a repository
duplicate            duplicate a model/dataset/Space under a new ID
move                 move between namespaces
branch / tag         Git-like reference management
delete-files         delete specific files by pattern

duplicate is especially useful. It is used when creating a derivative model based on someone else's repository. It is clearer than forking.

Specialized hub features


Datasets: hf datasets list, info, parquet
SQL:     hf datasets sql SQL โ€” run raw SQL with DuckDB against a dataset parquet URL
Models:  hf models list, info
Papers:  hf papers ls โ€” today's papers

hf datasets sql stands out. It throws SQL directly at a parquet URL. You can query a dataset without downloading it locally. This is fast when scanning a large dataset.

Discussions and PRs.


list, create, info, comment, close, reopen, rename
diff    view PR changes
merge   finalize a PR

It handles through the CLI the flow where people leave issues on a model repository you created.

Infrastructure and compute.


Endpoints: deploy, pause, resume, scale-to-zero, catalog
Jobs:     hf jobs uv (run Python with inline dependencies), stats (resource monitoring)
Spaces:   dev-mode, hot reload

scale-to-zero is practical. It is the method of paying only while you use an Inference Endpoint. It stays off the rest of the time.

Storage and automation.


Buckets:    create, cp, mv, rm, sync
Cache:      list, prune (remove detached revisions), verify (checksum)
Webhooks:   create, watch, enable/disable
Collections: add-item, update, list

cache prune is useful for disk cleanup. You do it later, but once you download several models the cache becomes tens of GB.

Global flags


--format json    machine-readable output
-q / --quiet    output only the ID

For an agent, --format json is almost mandatory. If you try to parse a human-readable log, the format breaks. And reducing output volume saves context.

There are extensions too.


Extensions: hf extensions install REPO_ID (extend CLI functionality from a GitHub repository)

Part 7 Summary

SkillVersionWhat it does
llama-cpp2.1.2GGUF quantization selection and local llama.cpp inference, URL-first discovery
serving-llms-vllm1.0.1High-throughput serving with PagedAttention and continuous batching
huggingface-hub1.0.1Search, download, upload models and datasets with the hf CLI

The three form a single line from local to deployment. Find it on the Hub (3), choose the quantization level and download (1), and if needed deploy with high throughput (2).

And there is one way this part differs from the others. The other six parts covered "how to do it well," and this part covered "what to choose." Choose the quantization level, choose whether it is vLLM or llama.cpp, and check whether the hardware supports it. There is no single right answer in technical choices. There is only the cost of a wrong choice.

The Full List of 21

We have now covered all 21. Looking back at the series and regrouping, it is this.


Part 1  Skill basics        hermes-agent-skill-authoring, plan, computer-use
Part 2  Coding delegation   delegate-coding, delegate-debugging, test-driven-development
Part 3  Debugging           systematic-debugging, python-debugpy, node-inspect-debugger
Part 4  Quality management  requesting-code-review, codebase-inspection, code-reference
Part 5  Research            deep-web-investigation, grounded-citations, research-paper-writing
Part 6  Documents/notes     notion, obsidian, nano-pdf
Part 7  Model operations    llama-cpp, serving-llms-vllm, huggingface-hub

Part 1 pinned down the definition of what a skill is, and Part 7 came all the way to how to choose a model. The five parts in between are all the hands-on work in between. Delegating, verifying, fixing, checking, citing, and handling documents.

One thing common to all of them shows here. These skills are all built to preempt the points where an agent causes an accident. The instructions inside the diff in Part 4, the uncited claim with no ledger in Part 5, the unshared page that spits out a 404 in Part 6. All of them are "points the agent passes through without knowing." Blocking those points one by one is what a skill is.