You know the difference between a junior leaning on Claude Code and someone who makes it hum. The junior types commands, gets something that works, and calls it done. The senior builds a skill — a reference file the agent can follow instead of reasoning from scratch each time. No more inconsistent output, no more watching the same model hallucinate different structures on the same request.
A skill is a markdown file (typically skill.mdcodecode) sitting inside your project’s .claude/skills/codecode directory. It opens with YAML front matter: the name, a one-line description that acts as the trigger, and tags. The description is what the agent scans when deciding whether to invoke this skill. Below that, the body: markdown instructions, optional script references, and a template for the output. The entire file is only loaded when the agent’s job matches that description — so you can stack hundreds of them without slowing a run.
The bottleneck shifts from knowing what to do to knowing which skills to compose and when to override them.
The problem most people hit: they write a skill that describes the outcome but fails to encode the process. The agent still guesses at half the steps. That’s where a structured methodology pays off. I’ve iterated through dozens of skills across Claude Code and its siblings, and six patterns consistently separate a skill that works from one that wastes tokens.
Step 1: Reverse-Engineer the Output
Start with a known-good result. Don’t write the skill first and hope the agent lands close. Execute the task manually — or get a previous version of the output you already approved — and feed that into the agent with a prompt: “Analyze this output. Reverse-engineer the steps required to produce it from the raw inputs.”
This forces the model to infer the implicit decision points. You’ll see it surface formatting rules, prioritization logic, and verification steps you never consciously wrote down. Capture those as the skill’s body. If you start from a vague goal (“write a research brief”), you get a vague skill. If you start from a concrete artifact, you get a concrete procedure.
Step 2: Atomic Tasks, Specific Triggers
A skill should own one task, not a department. “Manage the marketing pipeline” is too broad. “Write a weekly social media post from a YouTube transcript” is right. The YAML description must be tight enough that the agent knows exactly when to invoke it. I use a naming convention: research-briefcodecode, x-article-from-videocodecode, email-draft-in-stylecodecode. Each gets its own file.
Inside the body, break the procedure into numbered steps. Each step should produce a checkable artifact — a file, a summary, a screenshot — so the agent has a clear milestone. This granularity also makes it easier to chain skills later. You can call one skill to transcribe a video and pipe that output into another skill that formats it as a tweet thread.
Step 3: Calibrate the Freedom Level
Not all skills need the same degree of prescription. A deterministic task — data extraction, format conversion — benefits from rigid step-by-step instructions. “Open file X, read column Y, write to CSV with headers A, B, C.” No wiggle room.
A non-deterministic task — writing an article, summarizing a debate — needs guardrails, not scripts. Tell the agent what constraints apply (target audience, tone, source priority) and what a good output looks like, but leave room for judgment. Over-constraining a creative task makes the output generic; under-constraining a mechanical one invites drift.
The key insight: the same skill file can mix deterministic and non-deterministic sections. You might say “Step 4: Identify the top three sources. Use the following method for ranking. Step 5: Write a 300-word summary. Follow the style guide but adapt to the source material.” The model respects the boundary because you defined it.
Step 4: Build Verification Into the Skill
The most common failure mode: the skill produces output that looks plausible but fails a basic quality check. A skill without a verification step is a half-baked tool.
Add a section at the end: “After completing the output, run the following checks. For each check, either pass/fail or report the result.” Objective checks are easy: “Confirm every claim has a citation.” “Verify the word count is between 500 and 600.” Subjective checks require a prompt: “Read the article aloud. Does it flow logically? If not, rewrite the transitions.”
You can go further: have the same agent or a second agent review the output against the original skill instructions. This “LLM-as-judge” loop catches both formatting errors and logical gaps. I often include a boilerplate verification step: “Open the output file, render it, and take a screenshot. Confirm no visual artifacts.” The agent can use browser tools to do this.
Over time, you’ll update the verification section based on observed failures. Every time you notice the agent made a mistake, add a line to the pitfalls area. That list becomes the institutional memory of your skill.
Step 5: Down-Model Aggressively
Most people run every skill on the largest model they can access. That’s expensive and often unnecessary. A well-written skill — one that encodes the procedure precisely — can produce acceptable results on a cheaper model.
Build and debug the skill on a capable model (Claude Opus or GPT-4 class). Once it works consistently, drop down to a smaller model (Claude Haiku, GPT-4o-mini). Run the same test cases. If the output degrades — less structure, more hallucination — move up one tier. I’ve seen skills that require Opus for the verification loop but run the main generation on Haiku with no loss.
This isn’t one-size-fits-all. A skill heavy on browser interaction or dynamic code execution may need a larger model’s context window. But the principle holds: test the floor before settling on the ceiling.
Step 6: Treat Every Run as a Training Cycle
A skill is never “done.” Each time you run it, scan the output for improvements. Did it format the report correctly but miss the executive summary? Tell it: “Update the skill to always include an executive summary as the first section.” Then ask the agent to edit the skill file itself.
This “bike method” — analogous to teaching a kid to ride, with constant feedback and incremental adjustments — compounds over time. After ten iterations, the skill encodes not just the process but the edge cases you’ve encountered. It learns from its mistakes because you recorded them in the file.
I use a simple protocol: after each run, type “Feedback for skill: [what worked, what didn’t]. Update the skill file accordingly.” The agent reads its own output, identifies the gap, and rewrites the relevant section. Next run, the fix is permanent.
Where Skills Live and How to Deploy
In Claude Code, skills live under .claude/skills/codecode. Each skill is a subdirectory with a skill.mdcodecode file. You can also include a scripts/codecode folder for helper code and a templates/codecode folder for output skeletons. The agent loads them lazily — it only reads the description front matter of all skills until a job matches, then it opens the full file.
Deployment is as simple as copying the folder into your project. For shared teams, you can version-control the skills directory. Alternatively, Claude Code supports global skills stored in a user-level configuration path, so the same skills are available across all your projects.
To test a new skill, run a known input and compare the output to your manual baseline. I keep a test suite of three to five cases that cover the range of inputs the skill should handle. If the skill passes all of them on two different models, it’s ready for production.
Common Failure Modes
Description too vague. The agent doesn’t invoke the skill because the one-liner doesn’t match the job it parsed. Write the description as a beam — “Use this skill when the user asks to convert a YouTube video into an X (Twitter) article” — not a fuzzy category.
Instructions too loose. The agent follows the letter but not the intent. Tighten the procedural steps. If the output still varies, add a verification step that enforces a fixed structure.
No edge-case handling. The skill works on the happy path but breaks on empty inputs, conflicting sources, or missing data. The pitfalls section is the right place to document these. Each failure you catch becomes a guardrail.
Over-optimization too early. Don’t fine-tune the skill for the 5% edge case before it works consistently on the 95% common case. Ship the 80% version, run it for a week, then iterate.
Where the Field Is Going
The next evolution is skill composability — chaining multiple skills into a workflow that runs unsupervised. Claude Code now supports a runbookcodecode concept where you sequence skills with conditional branching. Instead of a single monolithic skill, you write small, testable modules and wire them together.
The teams that master this will build internal agent libraries that encode their entire standard operating procedures. New hires don’t need to learn tribal knowledge from a human; they invoke the same skills the senior team uses. The bottleneck shifts from knowing what to do to knowing which skills to compose and when to override them.
That’s a much more interesting problem — and one that no skill file can solve alone.
- What is a Claude Code skill?A skill is a markdown file (typically skill.md) inside your project's .claude/skills/ directory with YAML front matter and instructions.
- How do you start building a skill?Start by reverse-engineering a known-good output to infer the implicit decision points.
- What is the most common mistake when writing skills?Writing a skill that describes the outcome but fails to encode the process.
- How should skills be structured?Skills should be atomic tasks with specific triggers, each owning one task.
- What is the future of Claude Code skills?The next evolution is skill composability — chaining multiple skills into unsupervised workflows.