Automating Research with AI Agents: A Case Study in Skill Engineering

Discover how to build a reusable skill for the Hermes AI agent to automate weekly research briefs with consistent output.

By Central
This case study details the process of creating a skill file for the Hermes agent to automate research tasks.
Highlights
  • Verification loops matter more than the initial instructions when building a skill.
  • The agent can check its own work through objective and subjective verification loops.
  • For simple extraction tasks, the cheaper Luna model performed identically to Astra.

The problem is familiar to anyone who has relied on an LLM for regular research: the output changes shape every time. One run hands back a numbered list with citations; the next skips the numbers and buries the links. The model isn’t the culprit — inconsistent instructions are. Without a stable reference, the agent reinvents its approach on every request. That variability makes delegation impossible. You can’t trust it to run unsupervised.

This case study documents one real solution: building a reusable skill for an open-source agent called Hermes. The skill automates a weekly research brief — the kind of task that usually eats an hour of manual prompting and editing. The goal was to produce a consistent, citable document every weekday at 6 a.m. with zero human intervention.

Verification loops matter more than the initial instructions when building a skill.

The Problem: Inconsistent Output from the Same Prompt

Before the skill existed, asking the agent to research a topic produced a different result each time. The first attempt might include sources inline; the second would drop them. The structure would shift — each run re-planned the task from scratch. That’s fine for ad‑hoc questions, but for a recurring report, it means re‑reading every output to catch omissions.

The root cause wasn’t the model. It was the absence of a stored procedure. Every request carried the entire context in the prompt, but that context died when the session ended. No reference file told the agent what “done” looked like.

The Approach: Building a Research Skill with a Purpose-Built Tool

The fix requires three elements: an agent identity file, a machine to host it, and a skill file that defines the repeatable procedure. We used Hermes (MIT‑licensed, ~236k GitHub stars) on a Hostinger KVM‑2 plan with 2 cores and 8 GB RAM. The skill was authored using a free web tool called Ordain (ordain.host), which writes both the identity file and the skill file to the agent’s native format.

Step 1: Define the Agent’s Identity (soul.md)

The identity file lives at /opt/data/soul.mdcodecodecode. It sets tone, independence level, and operational constraints. For this project, the agent was instructed to work in British English, deliver the answer before its reasoning, and explicitly flag uncertainty rather than fill gaps. That last rule is critical for research work — it prevents hallucinated facts.

Step 2: Write the Skill File (skill.md)

The skill file lives in hermes/skills/research-brief/skill.mdcodecodecode. Its structure is:

  • YAML front matter: name, one‑line description, tags. The description is the only part the agent reads to decide whether to invoke the skill. Too vague and it never opens the file.
  • When to use: scope conditions.
  • Procedure: step‑by‑step instructions, including source order (official docs before change logs before forum threads).
  • Pitfalls: known mistakes, such as “do not fill missing data with plausible numbers”.
  • Verification: a section that tells the agent what “finished” looks like and how to check its own work.

The verification block is the most overlooked part. Without it, the agent can follow the steps but output whatever format it wants. We added a line requiring the exact printed format for every field. The first test ignored that; the second matched perfectly.

Step 3: Deploy and Test

The soul.md file was uploaded to Hermes via Telegram, then the agent wrote it to the correct path. The skill file was uploaded the same way. Both are text files — editing later means changing a line and re‑uploading.

The acid test: run the same research request twice and compare. Before the skill, the two answers would differ in structure and fact selection. After the skill, they matched headings, source order, and even the “not found” entries. The only variation was new data found on the second run — exactly what you want from a real research tool.

Results: Consistent Output, Autonomous Schedule

After the skill was field‑tested and adjusted, the agent was told to run the brief every weekday at 6 a.m. and message the result via Telegram. It set the schedule itself. The first automatic run produced a document that opened with a one‑line summary, followed by every change from the last 90 days with a link next to each number, then a section of source disagreements (both sides written out), and finally a list of anything it couldn’t confirm — with the URLs it checked.

The quality was high enough that the human reviewer only needed to skim for blind spots. No iteration was required.

Lessons Learned: Four Skills Worth Building First

Not every task benefits from a skill. The best candidates are jobs you’ve done more than three times in the same way. Force-fitting a variable process into fixed steps makes the agent worse. For research-based roles, four skills pay back fastest:

  1. Research brief – the one built here. It automates the most repetitive part of staying current.
  2. Inbox triage – drafts replies in your voice, leaving you to edit rather than compose from nothing.
  3. Content scripting – produces a script in your structure, not the model’s default.
  4. Outreach – finds candidates and writes the first message.

The key insight from building this skill is that verification loops matter more than the initial instructions. The agent can check its own work — both objective checks (source count, fact verification) and subjective ones (tone, flow). You train the skill by feeding it feedback after each run. Over time, the error rate drops and your trust rises.

The model used to run the skill can often be downgraded. Develop on Astra, then test on Soul, then Terra, then Luna. If the cheapest model produces the same quality, there’s no reason to pay for more compute. In our tests, Luna handled simple extraction tasks identically to Astra; only visual browser interactions required the higher‑end model.

The Future of Agentic Research

The research brief skill turned a weekly distraction into a background process. But the real shift is in how you think about your own workflow. Any task you can describe precisely enough to write down can be delegated. The limiting factor is no longer the technology — it’s your willingness to spend the first few cycles training the skill instead of doing the work yourself.

One caveat: skills are not set‑and‑forget. New tools, new data sources, and your own changing preferences will require updates. Treat each skill as a live document. The agent only learns what you explicitly teach it. If you stop giving feedback, it stops improving.

Share This Article