The 21 Latest AI Agent Skills, Part 5: Research with Sources and Writing Papers
Part 4 checked code. Part 5 is not code but words. Investigating, citing, and writing papers. Each of the three skills handles one of those stages.
They share something. None of them uses "what I know" as evidence. This is the strongest constraint in the entire series, and it shows up most strongly on the research side.
1. deep-web-investigation (v1.0.0)
The research category.
Conduct thorough, structured multi-source web research: iterative search,
source triangulation, cross-checking, and cited summaries.
Four words are the key. Iterative search, source triangulation, cross-checking, and cited summaries.
When to use it
- When a question is open-ended, fact-based, and time-sensitive ("what's the latest?", "compare these", "how does X actually work?")
- When a single search result is not enough and you need depth or verification
- When the user says "research", "investigate", "look into", "dig into"
The 9-step investigation procedure
This order is the entire skill.
1. Clarify the scope of the question (if ambiguous, the key terms, date range, depth)
2. Query plan: think of 3β6 different search angles. Vary the phrasing and use both Korean and English if needed
3. Run web_search for each angle. Refine the query iteratively based on results. Do not stop after one pass
4. Triangulate: choose 3β5 mutually independent and strong sources (official docs, trusted media, primary sources)
5. Extract promising pages with web_extract. Skim and keep key quotes and figures
6. Cross-check: confirm important claims with at least 2 independent sources. Record discrepancies too
7. Synthesize: organize by topic, not by source. Conclusion first
8. Cite: URL in the Sources: list. Mark claims from model knowledge not verified on the web with [unverified]
9. Report: structured output, concise and actionable. State confidence level and gaps
Step 4 states "prefer official/primary sources over blogs." Step 6 states "record discrepancies too." That is usually something people cannot do. Usually when you meet conflicting values, you pick one. Here you keep it.
Step 7 β "organize by topic, not by source" β is important and tricky. If you write the report in search order, it becomes a list of sources. That means nothing to a person.
Output format
## Investigation summary (Bottom line first)
2β4 sentences directly answering the question.
## Details
### <topic/criterion 1>
- Key fact (source name)
### <topic/criterion 2>
- ...
## Verification notes
- What agrees / what conflicts / unverified ([unverified])
Sources:
- [1] https://β¦ (source name)
- [2] https://β¦ (source name)
The third section is the characteristic of this skill. It attaches "what agrees and what does not" to the report. It is a section missing from most research deliverables.
Five rules
- Depth over speed: the default is 3 or more searches + 3 or more extractions
- No hallucination: assert only what extraction supports and mark uncertainty
- Do not overwrite the original: always synthesize, do not paste full text
- Respect timeliness: for "latest/current" questions, prefer recent sources and note dates
- Time cap: if the question is too big, do a focused pass, say what you covered, and suggest the next question
The last item is practical. If you get a broad question, you search endlessly and never finish. This skill deliberately chooses to stop and state how far it looked.
2. grounded-citations (v1.2.0)
The research category. Built jointly by Hermes Agent and Teknium.
Ground answers and documents in cited, verifiable sources.
If the previous skill is "find and organize," this one is closer to "do not get the citation numbers wrong." The version was bumped to 1.2.0, and this part was strengthened the most.
When to use it
The scope is defined broadly.
- Research, comparison, news summaries, "what is the current state of X?"
- Artifacts written to disk that cite, summarize, or report external facts β reports, briefs, documents, decks, wikis
- Fact-finding the user is likely to check later
- Multi-source synthesis where conflicting sources each need their own citation
Conversely, skip in these cases.
- When search is a side step of another task β a quick syntax/version check in the middle of coding, casual conversation, creative work
- Reveal URLs only when a link is genuinely worth using
The citation ledger
The core of this skill is sources.py. It runs on only the standard library. The ledger location is set per profile and can be changed with --ledger or an environment variable.
S=~/.hermes/skills/research/grounded-citations/scripts/sources.py
python "$S" reset # new ledger
python "$S" add https://example.com/a --title "A" # output: [1]
python "$S" add https://example.com/b --title "B" # output: [2]
python "$S" list # ledger table
python "$S" render # Sources: block
python "$S" verify draft.md # catch bad citations
add is idempotent and normalizes the URL. The same page always returns the same number within the ledger. No matter how many rounds of search and extraction you run, the numbers stay stable. This is the reason this skill exists.
Frequently used commands
reset clean ledger for a new task
add <url> [--title T] register a source and get a number
add <url1> <url2> ... several at once
ingest results.json register from JSON tool output
quote <id> --text "exact wording" --from page.txt
attach exact wording to a source
list [--json] view the ledger
render [--style markdown|plain|footnotes|bibtex|evidence] [--only 1,3]
render --cited-in draft.md only what the draft actually cited
render --replace-in draft.md rewrite the Sources block from the source of truth
verify draft.md [--strict] [--min-coverage 0.6] [--evidence]
render --cited-in is used most often in practice. If the ledger has 10 and the draft cites only 4, attaching the whole ledger instead of pulling out just those 4 makes the reader lose their way.
verify is the other side. It checks whether the draft's citations match the ledger and measures coverage. --min-coverage 0.6 means the minimum ratio of adjacent sentences. If a paragraph has no citation at all, it gets caught.
The pitfalls are half of this skill
If you transcribe the listed pitfalls as they are, the design intent of this skill is fully visible. They are all problems actually encountered.
- Registering after writing: the ledger must be filled from tool output. The moment you retrace it from the draft,
the hallucinated URLs the numbers were meant to eliminate come right back
- Changing numbers midway: do not edit IDs by hand. If the draft cites [4], [4] must be
that source. Reset only between tasks
- Do not retype URLs by hand into the Sources block: always render. A hand-typed URL is
an unverified claim
- Do not cite a page as if you had read it when you only read a search snippet: web_search's description only
supports what it said. If you need the body, start with web_extract
- Over-citing: 3 per sentence is the cap. Citing every clause makes it unreadable
- Do not attach citations to code/config artifacts: source comments go in the artifact body, not inside the code
- Parallel subagents: each subagent has its own working directory. They all must point at the same ledger so
numbers do not collide
- Making quote from a snippet: evidence quotes must come from the extracted page text. After web_extract, save
the file and use it with quote --from
- Paraphrasing with quote --text: it fails the exact-wording check. The answer is not to rewrite the sentence but
to find the actual sentence
- Do not use [unverified] as an escape hatch: attach it to claims that truly cannot find a source. If it is on most,
it means the search was insufficient
- Do not hand-edit the Sources block: use render --replace-in
I emphasize two in particular here. If you type the Sources block by hand, every one of those URLs becomes an unverified claim. And a search snippet and the actual page are different. A search result summary is just the first line of the document scrolled to, so if the content you want to cite was not in the snippet, it did not come from the source.
The skill warns especially strongly about [unverified]. It is the remaining problem the moment the label breaks out. If it is attached to most sentences, that means the investigation was insufficient β not that a label was missed. This is the pitfall research agents most often fall into.
3. research-paper-writing (v1.1.0)
The research category. Built by Orchestra Research. The target conferences are NeurIPS, ICML, and ICLR.
Write ML papers for NeurIPS/ICML/ICLR: designβsubmit.
Five philosophies
The second item here compresses what the previous two skills aim for.
1. Take initiative: do not just ask questions, produce a finished draft. Scientists are busy. Give something concrete and get a reaction
2. Never hallucinate citations: the error rate of AI-generated citations is about 40%. Always retrieve them programmatically.
Mark unverified citations as [CITATION NEEDED]
3. A paper is not a collection of experiments but a story: there must be one contribution stated in a single sentence.
If you cannot, the paper is not ready
4. Experiments exist for claims: every experiment must state which claim it supports
5. Commit early, commit often: commit with a message for every experiment batch and every draft update.
git log becomes the experiment history
The figure of about 40% is here. More than one in four citations written from memory does not exist. This is the direction the second skill entirely aims for, and the third skill confirms it with a number.
Initiative is split by level
High (repo is clear and contribution is clear) β write a finished draft, deliver, iterate on feedback
Medium (there is ambiguity) β draft with uncertainty marked, keep going
Low (there are major unknowns) β 1β2 targeted questions via clarify, then a draft
It is also split by section. Abstract, Introduction, Methods, Experiments, and Related Work all get "write it yourself" attached. This is not an option but a design. Do not stop at questions, but do not hand it over carelessly either.
Pin the contribution down first
Before writing anything in the paper, answer this.
- What: what is the one thing this paper contributes?
- Why: what evidence supports it?
- So What: why should the reader care?
And keep the TODO list as persistent state across sessions. A one-line contribution, literature review, experiment design, execution, analysis, draft, self-review (reviewer simulation), revision, submission prep.
The 5-step citation validation is mandatory
1. SEARCH β query specific keywords with Semantic Scholar or the Exa MCP
2. VERIFY β confirm the paper exists in 2 or more places: Semantic Scholar + arXiv/CrossRef
3. RETRIEVE β get the BibTeX programmatically via DOI content negotiation (no memory)
4. VALIDATE β confirm the claim you want to cite is actually in that paper
5. ADD β add the verified BibTeX to the references
If any step fails, mark it [CITATION NEEDED] and tell the scientist
The code to fetch BibTeX.
import requests
def doi_to_bibtex(doi: str) -> str:
response = requests.get(
f"https://doi.org/{doi}",
headers={"Accept": "application/x-bibtex"}
)
response.raise_for_status()
return response.text
A single "Accept": "application/x-bibtex" makes the DOI return BibTeX. That one line is the key. To summarize, VALIDATE in step 4 is essential. A paper existing and that paper supporting your claim are separate things. The case where you cite a paper that exists but says something different gets filtered out at this step.
A failed validation is left like this.
\cite{PLACEHOLDER_author2024_verify_this} % TODO: verify this citation
And the skill instructs you to always say this.
"I marked [X] citations as placeholders that need verification."
Do not hide it. If you quietly let it through, the person who discovers it later is the victim.
How to organize related work
You must group by method.
Good: "One line of work uses X's assumption [refs], but we use Y's assumption, and here is why..."
Bad: "Smith et al. introduced X. Jones et al. introduced Y. We combine the two."
Introducing papers one by one turns it into a table of contents. You have to write the relationships between methods for it to become a related-work section.
Part 5 Summary
| Skill | Version | What it does |
|---|---|---|
| deep-web-investigation | 1.0.0 | Iterative search, triangulation, cross-check, then topic-based synthesis |
| grounded-citations | 1.2.0 | Blocks URL errors and unverified citations with a citation-number ledger |
| research-paper-writing | 1.1.0 | Goes through 5-step citation validation to a paper draft |
The three form a single flow. Find it, number it, and shape it into a paper. All three skills force you not to use your own memory as evidence. That is the biggest accident an agent commits in research.
Next, Part 6 moves on to handing Notion, Obsidian, and PDF to the agent.
AI Knowledge Hub