All posts
AI Agents

AI Agent Evaluation: Unit-Test the Skills Your Agents Use

AI agent evaluation for skills: fixed test prompts, isolated runs, code checks plus a judge model, and a rerun after every prompt change.

October 3, 2026·7 min read·by Olexander Cheberko
Table of contentstap to expand

AI agent evaluation means running an agent on a fixed set of realistic tasks and scoring every run against criteria written in advance, so you can tell whether a change to its prompt, skills or model made it better or worse. For a skill, an eval works like a unit test: the same inputs every time, a pass or fail per check, and a rerun after every edit.

Below I apply the method Claude Code documents, and a loop that Lauren Tan, an engineer at Cursor, described in a recorded talk, to the agents that write this blog. It is a design, not a report, so this page has no eval scores.

What does an agent eval grade that an LLM eval does not?

The steps, not only the answer. A grader in Claude Code's plugin eval can check the final reply, the session transcript or a file Claude created (plugin evals), and the docs advise one grader on the result and one on the steps Claude took, such as which tool it called (choose graders). For a skill, seeing it trigger tells you Claude found it, not that it did what you intended, so Claude Code's docs measure invocation and output separately (evaluate a skill).

How do you unit-test a skill?

Against a baseline. Claude Code's docs say to collect a few realistic prompts and run each one in a fresh session with the skill available and again with it turned off, which for a personal or project skill means setting it to "off" in skillOverrides (evaluate a skill, skillOverrides). The session has to be fresh because leftover context from writing the skill masks gaps in its instructions. The skill-creator plugin automates the loop: it stores test cases in evals/evals.json, gives each case its own subagent with a clean context, writes pass or fail with evidence to grading.json, and runs a blind A/B between two versions of a skill so you can confirm an edit helps before you commit it (skill-creator).

How does Lauren Tan evaluate her skills?

In her talk, from about the 17-minute mark, Lauren Tan calls an eval a unit test for an agent and says you do not need a special framework to build one. Her coordinator agent writes a rubric for what the skill should do, then spawns sub-agents, each in its own directory, named so the sub-agent cannot tell it is being evaluated, because "agents can actually tell and when they do they change their behavior." The coordinator produces a score, and a judge agent on a different model can cross-check it so one model's bias does not decide the result. Then she hill-climbs: the agent loops on the eval until everything scores 10 out of 10, her example, and she reruns the eval every time she changes a skill.

How do you keep an agent from knowing it is tested?

Give the test run exactly what a real run gets and nothing that describes the test. Without --bare, claude -p loads the same context an interactive session would, including anything configured in the working directory or ~/.claude (headless mode), so a CLAUDE.md in the test folder that explains the rubric travels into the run, and a folder name that announces a test is exactly what Lauren Tan avoids. Claude Code's plugin eval handles this its own way: each run gets a temporary home and working directory, and a run cannot read the eval directory, so Claude never sees the graders or the sibling cases (how runs are isolated).

Which checks should be code, and which a judge?

Code wherever the rule can be written down. Claude Code's regex, tool_used, tool_order and file_exists graders are computed from the transcript and files and cost nothing, while an llm grader passes when a judge model votes PASS in at least two of three votes (grader types). A judge's answer can differ between runs, and more the longer the text it reads, so the docs suggest regex over a long generated file and the judge for short outputs, with rubrics written as concrete PASS and FAIL conditions (choose graders). --judge-model picks the judge, and the docs say to pin both the agent's model and the judge's in CI so scores stay comparable over time (run evals in CI).

What would an eval for the agent that writes this blog look like?

This blog's articles go through a multi-agent workflow: one writer agent per article, an independent fact-checker that re-reads the vendor's help pages and fixes or cuts claims, editor passes, and an agent that renders each page in a real browser at desktop and phone width. The writer's instructions have already changed after I rejected padded drafts, and only a fixed test set would show whether a change that fixes one article hurts the next. An eval for the writer would have four parts:

PartWhat it isWhere it comes from
Test topicsA fixed set that stays the same between runsSo scores before and after a prompt change compare like for like
Code checksscripts/content_lint.py: em dashes, links to removed or missing pages, over-long titles and descriptions, repeated FAQ questions, manuals whose checks differ from the declared steps, banned phrasingRules that already fail a draft today
Judge checksPASS and FAIL lines for what a regex cannot seeMy rejections of padded drafts and the fact-checkers' catches
RerunThe whole set after every prompt change, judged by a different model from the writerLauren Tan's routine in the talk

The judge lines I would start with:

PASS if the first two sentences answer the primary query.
FAIL if the draft opens with a story or a preview of what it will cover.
FAIL if a paragraph repeats a point an earlier paragraph made.
FAIL if a sentence says how often something happens without one real case behind it.
FAIL if a claim about a product has no link to that vendor's own documentation.

The first three come from early drafts I rejected as padded; the fix then was a fixed manual format and a memory rule, and an eval would show whether the rule survives the next prompt edit. The fourth comes from a fact-checker who caught a sentence that invented frequency, "the case I see most". The linter now bans that exact phrase, but a judge can catch a new wording of the same claim that no regex lists. The fifth scores the writer on what the fact-checkers do anyway: every product claim goes back to the vendor's own page. Following the docs' advice on long text, the judge would read excerpts, such as the opening and the FAQ block, rather than a whole draft. A Claude Code subagent can run on its own model through its model field (choose a model), which is how the judge would differ from the writer.

What will an eval not catch?

Claims about the world. Two of the fact-checkers' real catches would pass any rubric that reads only the text. One was a sentence about how Google Ads counts calls from ads that missed that, with call recording on, Google's AI can analyze the recording to judge lead quality (Google Ads Help). The other was an agent's finding that a site set consent to denied with no banner, which held only because the check ran from Bulgaria, and I withdrew it (testing Consent Mode from Europe). Catching those takes re-reading the vendor's page and rerunning a check from the right place, so the fact-checker stays in the workflow and I still approve every commit and push myself.

Where do evals stop and guardrails start?

An eval tells you whether a prompt change made the agent better on average; it does not stop one bad draft on the day it fails. That job belongs to hard checks that run on every file and block it, like the linter above. How I set those up is in AI agent guardrails as hard checks.

Tags

ai-agent-evaluationagent-evalsllm-evalsai-agent-testingclaude-code-skillsllm-as-a-judge

Frequently asked questions

What is the difference between agent evals and LLM evals?

An LLM eval scores one response to one prompt. An agent eval scores a whole run: the final result plus the steps the agent took to reach it, such as which tools it called and which files it wrote. Claude Code's eval docs suggest one grader on the result and one on the steps, so you learn both whether the answer was right and whether your skill produced it.

How many times should each agent eval case run?

More than once. Claude Code's plugin eval runs every case three times by default, because one run of a non-deterministic agent tells you little, and scores the case as the mean of its runs. Compare a prompt change on the same number of runs before and after.

What should an LLM judge grade in an agent eval?

Short outputs, against a rubric written as concrete PASS and FAIL conditions. A judge's verdict can differ between runs, and more so the longer the text it reads, so Claude Code's docs suggest grading a long generated file with regex checks and keeping the judge for short pieces.

Do agent evals cost money to run?

Yes. Every run of the agent under test and every judge call is a real model call, counted against your plan's usage or your API bill. A suite that also runs a no-plugin baseline makes about twice as many agent runs, so use free code checks where you can and keep judge calls for what code cannot decide.

Related posts