Building Research Skills for AI Agents: Most Tutorials Get It Wrong

Most tutorials for building AI research skills fail because they start with prompts instead of finished outputs. Here's how to fix it.

By Central
The article reveals two critical gaps in standard AI research skill tutorials and provides actionable fixes.
Highlights
  • Starting with a finished output instead of a prompt gives the agent a concrete example of good work.
  • Using one narrow task per skill with a specific trigger ensures the agent invokes the right skill.
  • For genuinely new topics, run the job manually first and document it before building a skill.

You read the guides. You wrote your skill file. The agent still handed you a mess. Every time.

The gap between what tutorials promise and what actually happens is not subtle. It is the difference between an agent that works the same way every week and one that invents a new bad approach every time you ask.

A skill for something you have never done is a skill for something you do not understand yet.

Here is where the standard advice fails — and what to do instead.

The Standard Advice (And Why It Crashes)

Most tutorials say three things. Write a clear description. List the steps. Let the agent figure it out.

That works for a one-off query. You are sitting there to catch the nonsense. But the moment you want that same research brief every Tuesday at 6 AM, the standard advice falls apart. The agent has no reference. It works out a new approach each run. No consistency. No trust.

The fix is not a better prompt. It is a better structure.

Gap 1: Starting with a Prompt Instead of a Finished Output

Standard advice says write a detailed prompt describing what you want.

Here is what that gets you. You write “research topics and give me a summary.” The agent returns a paragraph with no sources, no structure, and a fact that sounds right but is not. You rewrite the prompt. It returns something slightly different and still wrong.

What actually works: Start with the finished result.

Take a research brief you already made — one you are happy with. Hand it to the agent and say “reverse engineer this.” That completed brief is your definition of done. The agent can analyze what went into it, what sources you used, how you formatted it, and what you left out.

The creator of the six-step skill method calls this “engineering backwards.” You give the agent a concrete example of good instead of a description of good. Descriptions lie. Examples do not.

The standard advice assumes you can specify quality in abstract. You cannot. You can only show it.

Gap 2: Broad Skills Instead of Narrow Triggers

Standard advice says build one skill that handles all your research.

A single research skill covering “find information on any topic and summarize it” sounds efficient. It is not. The agent either opens it for everything (and slows down) or never matches the description (and you get generic answers).

What actually works: One very narrow task per skill.

Look at the YAML front matter in a skill file. It has a name, a description, and tags. That description is the only part the agent reads when deciding whether to open the rest. Write it too vaguely — “research tasksaa” — and the agent skips it on every job you built it for.

The fix: name exactly what triggers it. “research-brief-weekly” with a description that says “use when the user asks for a weekly research brief on AI regulation changes.” Nothing else.

In the source material, one effective skill is called “X article from YouTube video.” Not “content creation.” Not “social media posts.” One specific input, one specific output, one specific trigger. That is the pattern.

Gap 3: Ignoring the Deterministic vs Non-Deterministic Split

Standard advice treats all skills the same. Write steps. The agent follows them.

That fails because it ignores the type of automation you are building.

Deterministic skills — pull 500 records from a CSV, format them, put them in a table. Every input maps to the same output shape. These need rigid, exact steps. “Step one: open file. Step two: read row. Step three: write to output.”

Non-deterministic skills — write a research brief where each week’s topic is different, the sources change, and the conclusions are novel. These need constraints, not rigid steps. Tell the agent the sources to start from, the order to check them, and what to do when sources disagree. But leave the writing and selection open.

The mistake is treating a creative research brief like a data transfer. You over-constrain it and get generic output. Or you treat a data transfer like creative work and leave too much freedom, introducing errors.

The source material shows this clearly. The X article skill is highly non-deterministic. It does not say “put screenshot at line 400.” It says “capture relevant timestamps and place them where they support the argument.” The verification step then checks whether that happened. That is the right level of freedom.

Gap 4: No Built-In Verification

Standard advice says check the output yourself.

That defeats the purpose of an agent that runs while you sleep. If you must review every line, you saved no time.

What actually works: Build verification into the skill itself.

Every good skill has a verification loop. The agent produces the output. Then the same agent (or a sub-agent) checks it against criteria you defined. Objective checks — “every claim has a link” — and subjective checks — “does this read like the formatting in the example?”

The verification step is where you put the mistakes you have already watched the agent make. Write them down. “Do not assume a source is neutral because it ranks first. Check the domain.” Each run, the agent adds to that list. The skill improves.

One practitioner reports their X article skill iterates through verification cycles until it hits a quality threshold. Sometimes version 7. Sometimes version 2. The agent decides when it is done, not the human.

Without verification, a skill that reads perfectly can still hand back rubbish. Because nothing in it ever told the agent what “finished” looks like.

Gap 5: Using the Best Model for Everything

Standard advice says use the most capable model.

That is expensive and often wasted.

What actually works: Build with the best model, then reduce.

Develop the skill on the most capable model (Astra in the Codex ecosystem). Get it working. Then test it on each cheaper model down the stack — Soul, Terra, Luna. If the Terra produces the same quality at half the cost, that is your production model.

The source material demonstrates this live. A complex X article skill ran on four models. Luna — the cheapest — produced results that were subjectively better than Terra, a more expensive model. The cost savings were significant. The quality was not.

Not every skill can reduce. The YouTube-to-X skill needs better browser control (screenshots, cropping, formatting) than cheaper models provide. But many deterministic skills — data extraction, formatting, simple summaries — run well on the cheapest model.

Test. Do not assume.

Gap 6: The Illusion of “Finished”

Standard advice treats a skill as a one-time build.

Nothing could be more wrong.

What actually works: The bicycle method.

Every time you run the skill, give feedback. “This screenshot was poorly placed. The summary missed the main conflict. The formatting broke on mobile.” Tell the agent to update the skill file with that feedback.

Each run, the agent absorbs the correction. Next time, it does not make that mistake. Over weeks, the skill converges on your actual standard — not the standard you described, but the one you enforce.

The source material captures this precisely. The creator says they give feedback almost every single time they run a skill. The skill is never truly finished. It is always getting better because every execution teaches it something.

The bicycle analogy works. Start with training wheels (heavy oversight). Remove them as confidence grows. But never leave it unattended on a highway without some guardrails.

Standard vs. Practical: A Comparison

What Tutorials Say What Actually Works Why It Matters
Write a detailed prompt describing the task Reverse engineer from a finished example Examples define quality; descriptions guess at it
One skill handles all your research One narrow task per skill with a specific trigger Broad skills get skipped; narrow skills get invoked
List steps in order; the agent follows them Distinguish deterministic from non-deterministic automation Wrong level of freedom introduces errors or generic output
Check the output yourself Build verification loops into the skill itself Verification lets the agent run unsupervised
Use the most capable model Test cheaper models; deploy the cheapest that passes quality Unnecessary model cost adds up fast on scheduled runs
Build it once, run it forever Give feedback every execution; the skill improves continuously A skill that never changes gets worse as your standards evolve

The One Scenario Where Even This Advice Fails

All of this assumes the research topic fits a pattern you already know. You have seen the territory before. You can define what good looks like.

But what if the topic is genuinely new — a technology that emerged last week, a market that did not exist, a context with no prior examples?

In that scenario, reverse engineering from a finished output is impossible. There is no finished output. The verification criteria are guesses. The skill description is provisional.

The fix is not to build a skill at all. Run the job manually for the first few cycles. Document what you did. Then build the skill. The bicycle method requires at least one ride to know what training wheels look like.

A skill for something you have never done is a skill for something you do not understand yet.

Share This Article