{"id":96612,"date":"2026-10-03T03:01:00","date_gmt":"2026-10-03T07:01:00","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=96612"},"modified":"2026-09-26T09:49:08","modified_gmt":"2026-09-26T13:49:08","slug":"automating-research-ai-agents-96612","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/automating-research-ai-agents-96612\/","title":{"rendered":"Automating Research with AI Agents: A Case Study in Skill Engineering"},"content":{"rendered":"<p>The problem is familiar to anyone who has relied on an LLM for regular research: the output changes shape every time. One run hands back a numbered list with citations; the next skips the numbers and buries the links. The model isn\u2019t the culprit \u2014 inconsistent instructions are. Without a stable reference, the agent reinvents its approach on every request. That variability makes delegation impossible. You can\u2019t trust it to run unsupervised.<\/p>\n<p>This case study documents one real solution: building a reusable skill for an open-source agent called Hermes. The skill automates a weekly research brief \u2014 the kind of task that usually eats an hour of manual prompting and editing. The goal was to produce a consistent, citable document every weekday at 6 a.m. with zero human intervention.<\/p>\n<h2>The Problem: Inconsistent Output from the Same Prompt<\/h2>\n<p>Before the skill existed, asking the agent to research a topic produced a different result each time. The first attempt might include sources inline; the second would drop them. The structure would shift \u2014 each run re-planned the task from scratch. That\u2019s fine for ad\u2011hoc questions, but for a recurring report, it means re\u2011reading every output to catch omissions.<\/p>\n<p>The root cause wasn\u2019t the model. It was the absence of a stored procedure. Every request carried the entire context in the prompt, but that context died when the session ended. No reference file told the agent what \u201cdone\u201d looked like.<\/p>\n<h2>The Approach: Building a Research Skill with a Purpose-Built Tool<\/h2>\n<p>The fix requires three elements: an agent identity file, a machine to host it, and a skill file that defines the repeatable procedure. We used Hermes (MIT\u2011licensed, ~236k GitHub stars) on a Hostinger KVM\u20112 plan with 2 cores and 8 GB RAM. The skill was authored using a free web tool called <strong><a href=\"https:\/\/ordain.host\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Ordain<\/a><\/strong> (ordain.host), which writes both the identity file and the skill file to the agent\u2019s native format.<\/p>\n<h3>Step 1: Define the Agent\u2019s Identity (soul.md)<\/h3>\n<p>The identity file lives at <code>\/opt\/data\/soul.md<\/code>codecodecode. It sets tone, independence level, and operational constraints. For this project, the agent was instructed <a href=\"https:\/\/overcentral.com\/en\/best-places-to-work-awards-deadline-79738\/\" title=\"Best Places To Work Awards Extends Special Awards Deadline\" data-iacss-internal=\"1\">to work<\/a> in British English, deliver the answer before its reasoning, and explicitly flag uncertainty rather than fill gaps. That last rule is critical for research work \u2014 it prevents hallucinated facts.<\/p>\n<h3>Step 2: Write the Skill File (skill.md)<\/h3>\n<p>The skill file lives in <code>hermes\/skills\/research-brief\/skill.md<\/code>codecodecode. Its structure is:<\/p>\n<ul>\n<li><strong>YAML front matter<\/strong>: name, one\u2011line description, tags. The description is the only part the agent reads to decide whether to invoke the skill. Too vague and it never opens the file.<\/li>\n<li><strong>When to use<\/strong>: scope conditions.<\/li>\n<li><strong>Procedure<\/strong>: step\u2011by\u2011step instructions, including source order (official docs before change logs before forum threads).<\/li>\n<li><strong>Pitfalls<\/strong>: known mistakes, such as \u201cdo not fill missing data with plausible numbers\u201d.<\/li>\n<li><strong>Verification<\/strong>: a section that tells the agent what \u201cfinished\u201d looks like and how to check its own work.<\/li>\n<\/ul>\n<p>The verification block is the most overlooked part. Without it, the agent can follow the steps but output whatever format it wants. We added a line requiring the exact printed format for every field. The first test ignored that; the second matched perfectly.<\/p>\n<h3>Step 3: Deploy and Test<\/h3>\n<p>The soul.md file was uploaded to Hermes via Telegram, then the agent wrote it to the correct path. The skill file was uploaded the same way. Both are text files \u2014 editing later means changing a line and re\u2011uploading.<\/p>\n<p>The acid test: run the same research request twice and compare. Before the skill, the two answers would differ in structure and fact selection. After the skill, they matched headings, source order, and even the \u201cnot found\u201d entries. The only variation was new data found on the second run \u2014 exactly what you want from a real research tool.<\/p>\n<h2>Results: Consistent Output, Autonomous Schedule<\/h2>\n<p>After the skill was field\u2011tested and adjusted, the agent was told to run the brief every weekday at 6 a.m. and message the result via Telegram. It set the schedule itself. The first automatic run produced a document that opened with a one\u2011line summary, followed by every change from the last 90 days with a link next to each number, then a section of source disagreements (both sides written out), and finally a list of anything it couldn\u2019t confirm \u2014 with the URLs it checked.<\/p>\n<p>The quality was high enough that the human reviewer only needed to skim for blind spots. No iteration was required.<\/p>\n<h2>Lessons Learned: Four Skills Worth Building First<\/h2>\n<p>Not every task benefits from a skill. The best candidates are jobs you\u2019ve done <a href=\"https:\/\/overcentral.com\/en\/google-hollywood-ai-licensing-79386\/\" title=\"Google Needs Hollywood More Than Studios Need AI\" data-iacss-internal=\"1\">more than<\/a> three times in the same way. Force-fitting a variable process into fixed steps makes the agent worse. For research-based roles, four skills pay back fastest:<\/p>\n<ol>\n<li><strong>Research brief<\/strong> \u2013 the one built here. It automates the most repetitive part of staying current.<\/li>\n<li><strong>Inbox triage<\/strong> \u2013 drafts replies in your voice, leaving you to edit rather than compose from nothing.<\/li>\n<li><strong>Content scripting<\/strong> \u2013 produces a script in your structure, not the model\u2019s default.<\/li>\n<li><strong>Outreach<\/strong> \u2013 finds candidates and writes the first message.<\/li>\n<\/ol>\n<p>The key insight from building this skill is that verification loops matter more than the initial instructions. The agent can check its own work \u2014 both objective checks (source count, fact verification) and subjective ones (tone, flow). You train the skill by feeding it feedback after each run. Over time, the error rate drops and your trust rises.<\/p>\n<p>The model used to run the skill can often be downgraded. Develop on Astra, then test on Soul, then Terra, then Luna. If the cheapest model produces the same quality, there\u2019s no reason to pay for more compute. In our tests, Luna handled simple extraction tasks identically to Astra; only visual browser interactions required the higher\u2011end model.<\/p>\n<h2>The Future of Agentic Research<\/h2>\n<p>The research brief skill turned a weekly distraction into a background process. But the real shift is in how you think about your own workflow. Any task you can describe precisely enough to write down can be delegated. The limiting factor is no longer the technology \u2014 it\u2019s your willingness to spend the first few cycles training the skill instead of doing the work yourself.<\/p>\n<p>One caveat: skills are not set\u2011and\u2011forget. New tools, new data sources, and your own changing preferences will require updates. Treat each skill as a live document. The agent only learns what you explicitly teach it. If you stop giving feedback, it stops improving.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>The problem is familiar to anyone who has relied on an LLM for regular research: the output changes shape every time. One run hands back a numbered list with citations; the next skips the numbers and buries the links. The model isn\u2019t the culprit \u2014 inconsistent instructions are. Without a stable reference, the agent reinvents [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":98846,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/96612.png","fifu_image_alt":"Automating Research with AI Agents: A Case Study in Skill Engineering","footnotes":""},"categories":[31],"tags":[],"class_list":["post-96612","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology"],"fifu_image_url":"https:\/\/cards.overcentral.com\/cards\/en\/96612.png","fifu_image_alt":"Automating Research with AI Agents: A Case Study in Skill Engineering","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/96612","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=96612"}],"version-history":[{"count":1,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/96612\/revisions"}],"predecessor-version":[{"id":98847,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/96612\/revisions\/98847"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/98846"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=96612"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=96612"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=96612"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}