Evals¶
ai-rulez gives eval cases a supported place to live, keeps them out of the context your agents load, ships them in
plugin bundles when you ask, and can report skills that have none. It also defines a harness-neutral case format,
runs the cases through a pluggable runner (ai-rulez eval run), scores each skill (pass rate, trigger precision
and recall, ablation delta, token cost), records the scores next to the skill's content digest, and gates on them in
strict validation. It does not implement an agent: a runner does the running. A rubric is
graded by whatever the runner provides, or, with --grader builtin, by the model layer's judge.
- Case format
- Running evals
- Scores and the results file
- Linting cases and results
- Reports
- Running evals in CI
Layout¶
.ai-rulez/
skills/
deploy-staging/
SKILL.md
evals/ # cases for this skill
trigger-basic.eval.yaml
prompts/
long-request.md # referenced by prompt_file
fixtures/
notes.txt # referenced by files[].source
evals/ # optional: cases filed per skill name, away from the skill directory
deploy-staging/
handoff.eval.yaml
README.md
skills/<name>/evals/is a recognized directory.generatedoes not warn about it, and it is never written into any per-tool skill tree (.claude/skills,.agents/skills, ...): eval cases must not cost context..ai-rulez/evals/is an optional project-level tree. Use it when you prefer to keep cases away from the skill directory. Cases are filed per skill:.ai-rulez/evals/<skill-name>/counts as that skill's cases, and a case file outside such a directory belongs to no skill and is never run.- ai-rulez parses only files named
*.eval.yaml,*.eval.ymlor*.eval.json(the case format below). Everything else in those directories (another tool'scase.yaml,prompt.md, graders, fixtures) is yours: ai-rulez never parses it, but it is hashed into the skill's cases digest (so editing a fixture re-runs the skill) and copied byte for byte into plugin bundles..DS_Store-style junk,results/andnode_modules/are left out of the hash.
Case format¶
A case file is YAML or JSON. It holds a cases list, or one case written at the top level (its id then defaults to
the file name). The JSON schema is schema/eval-case.schema.json.
# .ai-rulez/skills/deploy-staging/evals/deploy.eval.yaml
schema_version: 1
cases:
- id: deploy-basic
description: The plain request fires the skill and names the cluster
prompt: Deploy the billing service to staging
expect_trigger: true
near_miss: # look similar, must NOT fire the skill
- Explain how our staging and production environments differ
- Roll back the last production deploy
files: # created in the working directory first
- path: services/billing/config.yaml
content: "env: staging"
- path: notes.txt
source: fixtures/notes.txt # copied from next to this file
assertions:
- type: contains
value: staging-eu
- type: not_contains
value: production
- type: regex
value: 'deployed \d+ services?'
- type: file_exists
path: deploy.log
- type: command_exit
command: grep -q ok deploy.log
exit_code: 0
rubric: The answer names the staging cluster and the rollout command it ran.
rubric_min_score: 0.8 # default 0.7; applies to rubric and rubric_items
model: haiku # overrides --model for this case
tags: [smoke, deploy]
- id: unrelated-question
prompt: What is the capital of France?
expect_trigger: false
| Field | Meaning |
|---|---|
id |
Lowercase letters, digits, ., _, -. Unique per skill. Required inside a cases list. |
prompt / prompt_file |
The user prompt, inline or a file next to the case file. Exactly one. |
expect_trigger |
Required. Whether the skill should fire for the prompt. |
near_miss |
Prompts that look like they should fire the skill but must not. Each becomes a derived negative case <id>.near-miss-<n> tagged near-miss. Only on expect_trigger: true cases. |
files |
Fixtures: path (relative to the working directory) with inline content or a source file. |
assertions |
Deterministic checks, below. |
rubric, rubric_min_score |
Judged by a model grader that the runner supplies, or by --grader builtin. |
rubric_items |
A weighted checklist instead of rubric (the two are mutually exclusive): a list of text and optional weight (finite, above 0, default 1), at most 50. The grader scores the share of the total weight the answer satisfies, a number from 0 to 1, which rubric_min_score (default 0.7) is compared with. A runner that has only a free-text rubric receives the checklist rendered as one (Score the answer against this weighted checklist (total weight 6): 1. (weight 3) ...), so rubric_items works with every runner. |
model, tags |
Per-case model override, free tags. |
Assertions:
type |
Fields | Holds when |
|---|---|---|
contains / not_contains |
value, optional path |
the final answer (or the file at path) does / does not contain the text |
regex |
value (RE2), optional path |
the answer (or file) matches |
file_exists |
path, exists (default true) |
the file is / is not in the working directory |
command_exit |
command, exit_code (default 0) |
the command, run through the shell in the working directory, exits with that status; needs --allow-exec because it executes text from a case file. It gets a scrubbed environment: PATH, HOME, the temp directories, the locale and CI, never your credentials |
Every path must be relative and stay inside its directory (no leading / or ~, no ..). prompt_file and
files[].source are also checked after symlinks are resolved: a symlink (to a file or to a directory) whose target
lies outside the eval directory is refused, and so is an eval directory that itself resolves outside the config
directory. files[].path names a file for the runner to create in the working directory; ai-rulez validates it is
relative and inside, and the runner must keep it there. A case with no
assertions and no rubric is a trigger-only case: it measures recall (positive) or precision (negative) and nothing else.
Unknown fields are errors, so a typo cannot silently disable an assertion. Problems are reported by
ai-rulez validate as AR996 eval-case-invalid with the file and line.
Running evals¶
ai-rulez eval run # every skill with cases, claude-plugin-eval runner
ai-rulez eval run deploy-staging --ablation # one skill, also run it without the skill
ai-rulez eval run --dry-run # what would run and roughly what it costs (--estimate is an alias)
ai-rulez eval run --mode activation --surface retrieval # does the right skill rank first? offline, free
ai-rulez eval run --runner command --runner-command ./my-runner.sh --harness codex
ai-rulez eval run --format junit --output-dir eval-report # eval-report/eval-report.xml
eval run [skill...] flags:
| Flag | Meaning |
|---|---|
--harness |
Harness the cases run against (default claude); recorded in the results. |
--runner, --runner-command |
claude-plugin-eval, command, claude-native or codex-native. The default is claude-plugin-eval for the claude harness and command when --runner-command is set; claude-native or codex-native for --surface native. |
--claude-bin, --runner-arg, --runs, --judge-model |
claude-plugin-eval only: the executable, extra arguments (repeatable, for example --runner-arg --trust-plugin), runs per case, grader model. |
--timeout |
Time limit for one skill with either runner (default 30m). When it ends the runner's whole process tree is killed (a process group on Unix); the output the runner may produce is capped (8 MiB for claude, 64 MiB for a command runner's answer). |
--model |
Model for the cases. |
--ablation |
Also run every case without the skill and report the delta. |
--dry-run, --estimate |
List what would run with an estimated cost range (low, expected, high). Calls no runner, writes nothing. --estimate is an alias. |
--mode, --surface, --scope, --description-from |
--mode activation measures only whether the right skill is chosen; see Activation mode. |
--format, --output-dir dir |
json, markdown (default) or junit; with --output-dir the report goes to <dir>/eval-report.<md\|json\|xml>. |
--max-cost USD |
Cost control, see below. |
--date, $AI_RULEZ_EVAL_DATE |
The date recorded in the results. The clock is never read, so equal inputs give an equal file. |
--changed-only, --base REF |
Only skills with files changed against REF (default HEAD; committed, uncommitted and untracked). |
--force |
Ignore the result cache. |
--threshold R |
Pass rate (0 to 1) a skill needs (default [lint.evals] min_pass_rate, else 1). 0 records scores without gating on them; a value outside 0-1 is rejected. |
--allow-exec |
Run command_exit assertions. |
--no-write, --results FILE |
Do not update, or use another, results file. |
--price-in, --price-out |
USD per million tokens for the estimate (default: the built-in price table). |
Every flag is validated before any runner is started: an unknown --format, a --max-cost, --price-in or
--price-out that is NaN, infinite or negative, or a --threshold outside 0-1 is an error that costs nothing. Each
skill's result is written to the results file (atomically) as soon as it finishes, so a crash or Ctrl-C keeps the
skills that already ran; the rest are reported as not run and the command exits non-zero.
Exit status: 0 when every selected skill passes, 2 when a skill is below its threshold, errored, was skipped over
budget, or has invalid cases, 1 for a failure to run at all.
The claude-plugin-eval runner¶
Builds a throwaway plugin containing the skill (without its evals/) and translates each case into the directory
format claude plugin eval reads: evals/<case>/prompt.md plus graders/*.md. A tool_used: Skill grader observes
whether the skill fired (inverted for expect_trigger: false); contains, not_contains, regex and file_exists
become regex and file_exists graders; rubric becomes an llm grader. It then runs
claude plugin eval <plugin dir> --json <file> --no-publish --threshold 0 --ablation with-without|none [--runs N] [--model M] [--judge-model M] [--max-cost-usd X]
and reads the per-run JSON back, taking a majority vote over a case's runs. Any result file, results directory or case
directory left by an earlier run is deleted first, and a run that times out (--timeout) or is interrupted is an
error: a partial result is never scored. A non-zero exit that still wrote a complete result is scored, and the cases it
missed count as errors. On Windows the whole process tree is killed through a Job Object. Cases the tool cannot express
(command_exit assertions, files fixtures) are reported as skipped, not failed, and left out of the score.
The adapter never adds --trust-plugin itself: pass --runner-arg --trust-plugin once you trust the plugin
directory. It was written against the help text and interview prompt of Claude Code 2.1.289; the JSON shape it reads
(cases[].arms.with|without[] with graders[], costUsd, error) was not checked against a live run, and an
unrecognized document is an error rather than a silent pass. claude plugin eval publishes its report to claude.ai by
default; the adapter always passes --no-publish. claude gets the same scrubbed environment as the native activation
runner (PATH, HOME, locale, CLAUDE_CONFIG_DIR, ANTHROPIC_* and CLAUDE_CODE_* authentication and provider
variables); cloud and forge credentials are not passed.
The command runner¶
--runner-command CMD runs CMD through the shell once per skill. It receives the request on standard input and
prints the response on standard output; standard error passes through. AI_RULEZ_EVAL_PROTOCOL=1 and
AI_RULEZ_EVAL_SKILL=<id> are in its environment.
{
"version": 1,
"skill": { "id": "deploy-staging", "dir": "/abs/.ai-rulez/skills/deploy-staging", "digest": "sha256:..." },
"harness": "codex", "model": "haiku", "ablation": true, "max_cost_usd": 2.5,
"cases": [
{ "id": "deploy-basic", "prompt": "Deploy the billing service to staging", "expect_trigger": true,
"assertions": [{ "type": "contains", "value": "staging-eu" }], "rubric": "...", "tags": ["smoke"] },
{ "id": "deploy-basic.near-miss-1", "prompt": "Explain how ...", "expect_trigger": false, "near_miss_of": "deploy-basic" }
]
}
Cases are self-contained: near misses are expanded and prompt_file and fixture source are inlined. The response:
{
"version": 1,
"results": [
{ "case": "deploy-basic", "arm": "with", "triggered": true, "output": "deployed 3 services to staging-eu",
"work_dir": "/tmp/run-1", "rubric_score": 0.9, "input_tokens": 5200, "output_tokens": 410, "cost_usd": 0.031 },
{ "case": "deploy-basic", "arm": "without", "output": "I cannot deploy from here", "cost_usd": 0.012 }
],
"cost_usd": 0.043
}
- A case with
rubric_itemsarrives withrubricset to the rendered checklist as well, so a runner that only knowsrubricgrades it correctly;rubric_itemsis sent too. triggeredis required for thewitharm. Thewithoutarm (only with--ablation) needs notriggered.- Give
output(andwork_dir, for file assertions) and ai-rulez grades the assertions itself; or givepassedto report your own verdict on the outcome checks, which wins over local grading.rubric_score(0-1) is required for cases with arubricunlesspassedis set. skipped: truewith areasonleaves a case out of the score;errorcounts the case as a failure.- Costs and tokens are optional. The protocol version must be
1; an unknown case or arm is an error.
The built-in grader¶
By default a rubric is graded by the runner (rubric_score, or its own passed verdict). --grader builtin grades it
with internal/llm's judge instead, from the answer the runner returned, so the grade does not depend on each
runner's own judge:
- Consent. Sending a transcript to a model is opt-in at every level:
--allow-llm, and, from the user config or the environment (never the repository),allow_network = trueand an[llm]model; see LLM access for the trust rule, the key and Gemini through liter-llm. Without all of that the run is refused before any case starts;--dry-runbuilds no client and sends nothing.--grader-max-cost(default $0.25) caps the grader's spend with the model layer's fail-closed budget; the spend also counts towards--max-costand is reported asgrader_cost_usd. - What is graded. Every case with a
rubricorrubric_items, in both arms (the ablation needs both), fromoutput. The judge returns a score in [0,1] with a one-line rationale at temperature 0, which replaces arubric_scorethe runner gave and is compared withrubric_min_score(default 0.7); the rationale is in the report (rubric_note). A checklist is graded as one call: the score is the share of the total weight satisfied. When the runner also gave its ownpassedverdict (for example from its assertion graders), the rubric grade must pass as well: the case passes only when both do. A result with nooutputhas nothing to grade and fails the rubric (with a warning). A failed judge call leaves that case ungraded, which scores as a failure. - Treated as data. The transcript goes to the judge between markers that carry a token derived from the request;
the prompt names the exact closing line, so a look-alike marker in the transcript cannot end the fence, and runs of
three angle brackets in the transcript are broken up. Text in the transcript that addresses the grader or asks for a
score is treated as an injection attempt and scores 0. The reply must be exactly one JSON object.
Secret-looking text in the rubric or transcript makes that case ungraded (nothing is sent); an oversized transcript
is cut to its head and tail (64 KiB). The judge call gets at least 2,048 completion tokens (
llm.DefaultJudgeCompletionTokens, the floor of the model layer's own judge): a reasoning model such asgemini-2.5-flashspends its thinking out of that budget before the verdict, so a smaller one leaves a cut-off reply. - Choose the model deliberately. The grade is only as good as the judge. In a live comparison against Claude
Code's own rubric grading on 24 real transcripts,
gemini-2.5-flash-liteagreed on 79% (Cohen's kappa 0.60) and erred lenient (it passed a control rubric the answer plainly contradicted, and passed two answers that omitted a required detail), whilegemini-2.5-flashagreed on 96% (kappa 0.92). Check a cheap judge against a few labeled transcripts before trusting its pass rate. - Runners. The command runner must return
output.claude-plugin-evalreturns the answer its own llm grader read; with--grader builtinthat tool still runs its grader (so the rubric is judged twice and the tool's verdict is ignored) and the built-in grade decides. The claude adapter's results carry no transcript for cases without a rubric. - Caching. The grader's model and the judge prompt version are part of the cache key, so grading differently re-runs the skill.
Caching¶
A skill is not re-run when the results file already holds a run with the same cache key: the skill's sha256 digest,
the digest of its eval material (cases, fixtures, rubrics and graders), the runner and its own settings (the
--runner-command text and the content hash of its first word when that is a file, so editing the script re-runs; or --claude-bin, --runs, --judge-model and --runner-arg for claude-plugin-eval),
harness, model, the ablation setting, --allow-exec (it changes how command_exit assertions grade) and the
ai-rulez version. Editing a case or the skill, or changing any of those, re-runs it. --force ignores the cache.
Only real grading results are cached: a run in which any case errored (rate limit, missing credentials, no result
from the runner) or nothing was scored is not stored and is retried on the next run, so an outage never turns into a
cached failure. A failed grade (a case that ran and did not pass) is a result and is cached. The pass/fail verdict is recomputed from the stored
score against the current threshold, so lowering --threshold needs no re-run.
Records are signed per user. eval-results.json is a committed file, and the cache key is an unkeyed hash anyone
can compute, so a pull request could add a record that claims a pass. Each record therefore carries a mac
(HMAC-SHA256 of the record) made with a per-user key, eval-results.key in the user config directory ($XDG_CONFIG_HOME/ai-rulez/, else
~/.config/ai-rulez/; 32 random bytes, mode 0600, never in the repository). A record without a valid mac for your
key (committed from another machine, hand-edited, or when no key can be stored) is unverified: it is still shown
and linted, but it never satisfies a cache hit. The skill re-runs, the run reports a warning (the stored result is
unverified ... re-running), and the fresh record is signed. A record is also required to match the skill and cases
digests on disk. Consequence: results committed by CI or a teammate re-run once on each machine; a CI job that
restores eval-results.json from its own cache replays only if it keeps the same key file. Treat the key like
llm-cache.key: XDG_CONFIG_HOME set from an untrusted .envrc relocates it. A CI gate should still use
eval run --force, or verify the provenance of eval-results.json, since anyone who can edit the file can also
delete the mac.
Cost controls¶
--dry-run(alias--estimate) prints the number of agent runs, estimated tokens and USD as a range. The estimate is deterministic and offline: the harness's own overhead (2,000 tokens), the prompt, fixtures, the skill'sSKILL.md(counted with the embeddedcl100k_basetokenizer, an approximation), 600 output tokens per run, an extra grader call per rubric, times the runs per case (3 forclaude-plugin-evalunless--runs; the same number is passed to claude with--runs, so the estimate and the run agree), times two arms with--ablation. That is the expected figure. The low figure assumes 0.8 times the input tokens and 0.5 times the output tokens. The high figure assumes 1.3 times the input tokens plus one more pass over everything beyond the harness overhead (a tool loop that re-reads the skill and fixtures) and 3 times the output tokens. The multipliers are fixed defaults, pinned by tests; treat the whole range as an order of magnitude. Prices come from one built-in table shared with the model layer ([llm]price overrides do not apply to evals; use--price-inand--price-out). Haiku, sonnet and opus are listed under their short names too; no--modelis priced as sonnet, and the report says so (priced_as, and a line in the markdown). A model the table does not list is priced as sonnet for the estimate (the report says so), and with--max-costit is refused unless--price-inand--price-outare given.- The assumptions above are defaults.
[lint.evals.estimate]overrides them (overhead_tokens,assumed_output_tokens,activation_output_tokens,tool_loop_factor, andprice_in_per_mtokandprice_out_per_mtokfor what the runs are really billed at; zero keeps the built-in value, and--price-inand--price-outwin), andeval calibrate-estimateproposes measured values. The harness overhead matters most: Claude Code's own system prompt, tools and skill listing make a one-turn activation run cost about 24,000 input tokens, not 2,000. - An activation estimate (
--mode activation --surface native) is tighter, since a run is one turn with no fixtures and no tool loop: the overhead, the names and descriptions of the installed set and the prompt in, and 150 tokens out, times prompts times--runs; low is 0.9 times the input and 0.5 times the output, high 1.25 and 2. - Every run records the estimate next to what the runner reported, in the report (
estimate_vs_actual) and in the skill's record ineval-results.json(estimate:low_usd,expected_usd,high_usd,actual_usd,error= actual/expected - 1,in_range, and the run count, token split and assumptions the calibration reads). A runner that reports no cost leaves the actual out. The history shows how far off the estimate tends to be; it never leaves the machine. --max-cost USDis an advisory budget for the whole run (all skills together), not a per-skill cap and not a hard limit. It refuses to start when the estimate exceeds it (the expected figure for case runs, the high figure for--mode activation;--max-cost-mode expected|highchooses), hands each runner the remaining budget (max_cost_usd;claude-plugin-evalpasses it as--max-cost-usd), and skips the remaining skills once the reported spend reaches it (statusskipped-over-budget, exit 2). It is checked between skills; inside one skill only the runner can enforce it, and--timeoutbounds the time. When a runner reports more than the budget it was given, the report carries a warning. A runner that reports no cost and no tokens at all is assumed to have spent the whole remaining budget (with a warning), so the run stops instead of continuing on an unknown spend. Spend is counted conservatively: the larger of the sum of per-case costs and the runner's own total, tokens priced with--price-in/--price-outwhen no cost is reported, and a runner call that fails is assumed to have spent the whole remaining budget, so the run stops. Acommandrunner should enforcemax_cost_usditself: it is the only one that knows what it spends. Negative or non-finite costs and token counts in a response are rejected.--changed-only, the cache and--runskeep the number of runs down;tagsand skill names narrow it by hand.
Calibrating the estimate¶
ai-rulez eval calibrate-estimate # propose assumptions from the recorded runs
ai-rulez eval calibrate-estimate --model haiku --format json
Reads eval-results.json and fits the assumptions to what the runs reported: for each harness, model and kind of
run (case runs and activation runs are separate groups; different models are never mixed) the harness overhead moves
by the median gap between the input tokens the runner reported and the ones the estimate expected, per agent run, and
the output assumption (assumed_output_tokens, or activation_output_tokens for activation) likewise. The
tool-loop factor is left alone: one record cannot separate it from the overhead. The output shows the current and
proposed values, the median token error before and after, the recorded cost error (median and p90), and a
[lint.evals.estimate] table to copy into your configuration:
activation, harness claude, model haiku: 6 run(s)
overhead_tokens 2000 -> 23875
activation_output_tokens 150 -> 307
token error (median) +718% -> -0%; recorded cost error median 161%, p90 204%
input price per MTok $1 list -> $0.3209 billed (prompt caching)
(Real output from one native activation run of six skills with Claude Code and haiku: 120 agent runs. With the
proposed values, the next estimate of the same run was $1.14 against an actual $1.10, every skill within its
[low, high] range, instead of $0.44 against $1.19.)
The cost can be far below the token count priced at list: Claude Code caches its prompt, so most input tokens are
billed at a fraction of the list price. When the billed cost implies an input price more than 10% off the list price
the proposal adds price_in_per_mtok (the median over the runs, with output at the list price).
Only records signed with your key count; a run whose runner reported no token split is left out, and a group with
fewer than 3 runs (--min-samples) is marked low-confidence. The store keeps one record per skill and kind, so the
samples are the latest run of each skill. Nothing is written or sent anywhere; the command is deterministic and
offline.
Activation mode¶
A full case run answers "does the skill do the job". Activation mode answers a cheaper question first: "for these prompts, is the right skill the one that gets chosen, and the wrong one not?"
ai-rulez eval run --mode activation --surface retrieval
ai-rulez eval run deploy-staging --mode activation --surface retrieval --scope all --format json
The retrieval surface ranks every prompt of a skill's cases (including the expanded near_miss prompts) against
the competing skills with the same offline ranker the served-skills find_skill tool uses (BM25F over name, triggers,
keywords and description; see ai-rulez search). It calls no model and no network, costs nothing and
is deterministic. A skill fires for a prompt when it ranks first, so a positive prompt passes when it fires and a
negative one (expect_trigger: false, near misses) when it does not. Only the prompt and expect_trigger are used:
fixtures, assertions and rubrics are ignored (counted as ignored).
It measures the finder, not a model's own choice; the report says so. Per skill it reports:
- Recall, precision and the false activation rate, each with a 95% Wilson interval (small samples are the norm, and a bare percentage over four prompts says little).
- recall@1, recall@3 and MRR over the positive prompts, and the rank of the skill for every prompt.
- Stolen by: the siblings that ranked first on its positive prompts, with counts and shares, and a top-level
confusion matrix (
confusion[expected][won],nonewhen nothing ranked).
--scope domain (default) makes the skill compete with its own domain's skills plus the root skills; --scope all
with every skill. A skill alone in its scope gets a warning, since a stolen trigger cannot be measured. The report
also records a digest of the competing set: a sibling's edited description changes the competition and so the digest.
A skill passes when the share of passing prompts reaches --threshold (default 1). Exit status is as for a case
run: 2 when a skill fails, has invalid cases, or errors.
Comparing descriptions¶
ai-rulez eval run <skill> --mode activation --surface retrieval --description-from candidate.txt measures a
candidate description of that one skill on the same prompts: the file's text (at most 16 KiB, UTF-8, no hidden
characters) replaces the description in a scratch copy of the skill, on either surface, and the report carries a
warning saying so. The source is not edited and nothing is recorded in eval-results.json, so the run can be repeated
with another file and the two reports compared. It needs exactly one skill.
{
"schema_version": 1, "mode": "activation", "surface": "retrieval", "scope": "domain",
"skills": [{
"id": "deploy-staging", "status": "ran", "passing": false,
"recall": { "value": 0.5, "n": 2, "interval": { "low": 0.09, "high": 0.91 } },
"stolen_by": [{ "skill": "release-notes", "prompts": 1, "share": 0.5 }]
}],
"confusion": { "deploy-staging": { "deploy-staging": 1, "release-notes": 1 } }
}
The JSON follows schema/eval-activation.v1.schema.json; markdown is the default format, junit is
refused. Each measured skill's rates and ids (never prompts) are recorded in eval-results.json under activation,
signed like the rest of the record, and judged by AR9A1 and AR9A2 (off until you set a threshold). A later
eval run keeps the block; an activation run does not touch the case-run result. --dry-run/--estimate ranks and
prints but writes nothing; it still exits 2 when a skill fails its threshold, as a real run does.
The native surface¶
ai-rulez eval run --mode activation --surface native --model haiku --runs 5
ai-rulez eval run deploy-staging --mode activation --surface native --dry-run # the estimate, no model call
native asks a harness's model. For each skill it installs every competing skill (the --scope set) in the
harness at once, repeats each prompt of the skill's cases --runs times (default 5), stops each run at the first turn
and records which skills loaded. A one-skill harness can never show a stolen trigger; this one can. A prompt's
activation rate is the share of runs in which the skill under test loaded, reported with its Wilson interval and
the per-skill counts (fired_counts, with none for runs in which no skill loaded). A positive prompt passes at a
rate of at least 0.8, a negative one at most 0.2; a prompt that passes but would fail with one run going the other way
is borderline (3 or more runs, never a perfect score), so 4 of 5 reads differently from 5 of 5. Borderline prompts
count as passing. Recall, precision and false activation are computed from the thresholded prompts (as on the
retrieval surface); run_recall and run_false_activation pool the individual runs, with their own intervals. A
positive prompt that fails and that a sibling won most often counts towards that sibling's stolen_by. The confusion
matrix counts runs, not prompts. A prompt the runner gave no result for is an error prompt, the skill is reported
with an error and not scored, and nothing is recorded or cached for it.
A skill's measurement is stored in eval-results.json (activation: rates, the runner, harness, model, run count,
the cache key and the estimate next to what the run cost) and replayed, as cached, while the skill, the competing
set (a sibling's edited description changes it), the cases, the runner and its settings, the model, the run count and
the ai-rulez version are unchanged. --force repeats it. A stored record that is not signed with your key is never
replayed.
Runners and the capability handshake. A runner declares what it supports, and ai-rulez asks before it sends
anything: --surface native is refused (exit 1, "runner ... does not support activation mode") for a runner that does
not declare the activation capability and the native surface, never run as full cases. claude-native and
codex-native declare both. The command runner is probed: ai-rulez starts the command once with the request
{"version":1,"mode":"capabilities"} (no skill, no cases; AI_RULEZ_EVAL_MODE=capabilities in the environment) and
reads
A command that answers without capabilities (every runner written before this) is refused; one that answers the
probe with results is refused too, since it ran something. --dry-run skips the probe and every runner call.
The activation request a capable command receives:
{
"version": 1, "mode": "activation", "surface": "native", "harness": "claude", "model": "haiku",
"runs": 5, "max_turns": 1, "max_cost_usd": 0.5,
"skill": { "id": "deploy-staging", "dir": "...", "digest": "sha256:...", "description": "..." },
"skills": [
{ "id": "deploy-staging", "dir": "...", "digest": "sha256:...", "description": "Deploy a service to staging" },
{ "id": "release-notes", "dir": "...", "digest": "sha256:...", "description": "Write release notes" }
],
"cases": [
{ "id": "deploy-basic", "prompt": "Deploy the billing service to staging", "expect_trigger": true, "target": "deploy-staging", "runs": 5 }
]
}
skills is the installed set (the skill under test is in it); fixtures, assertions and rubrics are not sent. The
response has one with result per case and no without arm:
{ "version": 1, "results": [
{ "case": "deploy-basic", "arm": "with", "runs": 5, "fired_counts": { "deploy-staging": 4, "release-notes": 1 },
"input_tokens": 20500, "output_tokens": 600, "cost_usd": 0.0071 } ] }
runs (at least 1) and fired_counts (ids from skills, plus none; each count at most runs) are validated; a
result with error or skipped needs neither. errored_runs counts repetitions that failed and are not in runs; when more runs errored than completed, the
prompt has no result and the skill is not scored. input_tokens, output_tokens and cost_usd feed the estimate record.
claude-native(harnessclaude, the default for--surface native): writes every skill of the set into one throwaway plugin and runs, for each prompt and repetition,claude -p --output-format stream-json --verbose --max-turns 1 --no-session-persistence --setting-sources "" --permission-mode dontAsk --tools Skill --allowedTools Skill --plugin-dir <plugin> [--model M]in an empty directory, with the prompt on stdin. ASkilltool call in the transcript (ai-rulez-activation:<id>) is a load; cost and tokens come from the closingresultline (cache reads and writes count as input). Four runs are in flight at a time and the adapter enforces--max-costitself: once the spend reaches the budget no further run starts. No user, project or local settings are loaded, so the skills in your own Claude Code configuration do not compete; the harness's built-in skills and any enabled plugins still do, and Claude Code's own system prompt makes each run cost far more input than the default estimate assumes (seeeval calibrate-estimate). Run it with--model haikuunless you mean to pay for a larger model. The plugin holds only each skill'sSKILL.mdreduced tonameanddescription:hooks,allowed-tools, scripts and references are not installed, and a skill with an error-level security finding is refused before any run. Theclaudeprocess gets a scrubbed environment (PATH,HOME, locale,CLAUDE_CONFIG_DIR,ANTHROPIC_*andCLAUDE_CODE_*authentication and provider variables); cloud and forge credentials are not passed.codex-native(harnesscodex): writes the set to<work>/.agents/skills/<id>/and runscodex exec --json --ephemeral --ignore-user-config --ignore-rules -s read-onlyfrom stdin. Codex loads a skill by reading itsSKILL.md, which shows up as a shell command; the run is killed as soon as one is read or after three commands, so the model never acts on the task. The commands that did start run read-only in an empty directory withHOMEpointed at an empty directory and a scrubbed environment (CODEX_HOMEstays, for the login). Codex reports usage only when a run finishes, so this adapter reports no tokens and no cost.--max-costcannot be enforced through it, so a capped run (without--dry-run) is refused; run--dry-runfor the estimate, then run uncapped. It is experimental: verified live against codex-cli 0.160, one prompt per run costs on the order of 200,000 input tokens.--runner-command: any other harness; implement the probe and the request above.
--max-cost is checked against the high estimate figure before a native run starts (--max-cost-mode expected
checks the expected figure instead; for case runs the default is expected). Between skills, spend that reached the
cap skips the rest, as for case runs. --timeout bounds one skill's runner call.
Design decisions¶
--surfacehas no default: silently choosing the offline ranker would look like measuring a model, and choosingnativewould spend money.- Retrieval is one deterministic run per prompt, so a rate is 0 or 1 and a prompt is never borderline. The native thresholds (a positive passes at an activation rate of 0.8 or more, a negative at 0.2 or less) apply to repeated runs.
- "Borderline" is one run from failing, not "the Wilson interval straddles the threshold": at five runs every interval straddles 0.8, so that rule would mark everything.
- Repetition matters. In a live run of 24 prompts with Claude Code and
haiku, the same prompt fired its skill in 5 of 5 runs one time and 3 of 5 the next, and on 23 of 24 prompts the pass or fail verdict held between two repetitions. The offline ranker agreed with the model's behaviour on only 16 of those 24 prompts: it ranks the right skill first, while the model often loads a generic sibling (repository-layout) instead, so retrieval is a cheap pre-check, not a substitute fornative. - One turn (
max_turns1) is enough: the decision to load a skill is in the first turn. The design's default of 2 would pay for a second turn that does not change the answer. - The default scope for large catalogs stays
domain: bounded, at the price of missing a cross-domain steal; use--scope allto look for those. - The confusion matrix is reported once at the top level; each skill carries its own
stolen_bylist. - A record measured on an older skill digest, or one that is not signed with your key, is not judged by
AR9A1andAR9A2(an unsigned one is reported as unverified, likeAR997). - Not implemented:
--description-from(a candidate description for one run, from the design's phase 5).
Importing scenarios¶
ai-rulez eval import --from tessl ./scenarios/add-health-endpoint --skill http-service
ai-rulez eval import --from tessl ./scenarios --output-dir ./tmp-cases --dry-run --report import.json
eval import turns scenarios written for another tool into case files. It is offline: it reads local files only,
never contacts a service or needs its token, runs nothing it reads, and treats the input as untrusted text. Only
--from tessl exists; the code is a Source interface (Detect, Load, Map) so another format can be added.
A Tessl scenario is a task (task.md) plus a weighted checklist (criteria.json). The shape of criteria.json is
not verified: the importer follows public notes (a task, a weighted checklist, a pass percentage), not a schema or a
sample of the service's real files, so the mapping is tolerant (a few spellings of each field are read) and anything
it does not read is reported instead of guessed at. The documented shape, which is illustrative:
{
"scenario": "add-health-endpoint",
"criteria": [
{ "name": "adds route", "description": "Registers GET /health on the router", "weight": 3 },
{ "name": "returns 200", "description": "Handler returns status 200 with a JSON body", "weight": 2 },
{ "name": "no new dependency", "description": "Does not add a dependency", "weight": 1 }
],
"pass_threshold": 0.7
}
| Input | Becomes | Notes |
|---|---|---|
task.md |
prompt_file: <id>.task.md |
exact; the file is written beside the case |
scenario name (scenario, name, id, title, else the directory) |
case id, slugged to [a-z0-9._-] |
a repeated name gets -2, -3 |
| (the scenario targets a skill) | expect_trigger: true, tag imported:tessl |
an assumption, reported |
criteria (a list of objects or strings, or an object) |
rubric with a numbered, weighted checklist (--rubric-mode single, the default), or rubric_items with the weights kept (--rubric-mode items) |
name, description and weight are read as name/title/id, description/criterion/text/check/prompt, weight/points/score/max_score; no weights means every criterion weighs 1 |
pass_threshold (also passing_score, pass_score, passing_threshold, threshold) |
rubric_min_score |
a fraction (0-1), or a percent (above 1 up to 100, or "70%"); the unit it was read as is reported; none means the default 0.7, noted |
files, fixtures, starting_files, setup_files |
files |
inline content, or a source copied to fixtures/<id>/... beside the case; paths are checked like case paths |
activation block with should_not_trigger prompts |
near_miss |
reported |
anything else (baseline, repeats, agent, model, unknown fields) |
not imported | listed as unmapped (AR9A5, informational) with its JSON path, a short value and, where there is one, the flag it belongs to (--ablation, --runs, --model) |
A criterion with weight 0 is dropped (and reported as unmapped); a negative weight, a checklist whose weights sum to 0, a missing task, or a threshold above 100 is an error.
--lift-assertions (off by default) also converts criteria that state a mechanical check in a fixed phrasing, with the
path or the text quoted, into assertions: The file "x" exists / is created / is present, The file "x" does not
exist, The output|answer|response contains|includes|mentions "y" and ... does not contain "y". It is
conservative (an unquoted value, two quoted values or a path that is absolute or leaves the directory is not lifted),
never removes the criterion from the rubric, puts a # lifted from criterion "<name>" comment on each assertion and
lists every lift in the report. It is off by default because a wrong lift silently changes what is graded.
Output goes to .ai-rulez/skills/<skill>/evals/ (--skill) or --output-dir: <id>.eval.yaml with a provenance header
(importer version and a sha256 of the input files), the task, and any fixture copies. The command prints what was
mapped, assumed, lifted and left unmapped (--format json prints the same as JSON; --report FILE also writes it):
$ ai-rulez eval import --from tessl ./scenarios/add-health-endpoint --skill http-service
wrote add-health-endpoint.eval.yaml, add-health-endpoint.task.md (1 case)
mapped: task -> prompt_file
mapped: 3 criteria -> rubric (weights kept as text)
mapped: pass_threshold -> rubric_min_score 0.7 (read as a fraction)
assumed: expect_trigger: true (the scenario targets a skill)
Safety. Nothing is written unless every scenario maps, and an existing file is not overwritten without --force
(a symlink at a target is replaced, never written through). Input files are capped at 2 MiB and JSON at 32 levels,
must be UTF-8, and must be regular files (a symlinked task.md or fixture that resolves outside the scenario is
refused). Fixture and file paths follow the case format's rules (relative, no ..). The text of the task, the criteria,
fixtures and near-miss prompts is scanned with the security rules: hidden characters (AR002) and credentials
(AR001) refuse the scenario (the credential is masked in the message); an instruction-override phrase (AR004) is
flagged in the report as a warning, since a criterion ends up in a rubric a grader model reads. JSON keys that are not
plain identifiers are shown quoted in the report, so a hostile key cannot reach a terminal raw. Every file written
loads through the normal case parser (the AR996 check) before it is written.
Scores and the results file¶
Per skill, over the run's cases (near misses included, skipped cases excluded):
| Score | Definition |
|---|---|
pass_rate |
passed cases / scored cases. A case passes when the skill fired exactly as expect_trigger says and its assertions and rubric hold. Runner errors count as failures. |
trigger_precision |
TP / (TP + FP) over the with runs, where TP is a positive case that fired and FP is a negative case (near misses included) that fired. null with no denominator. |
trigger_recall |
TP / (TP + FN). null with no denominator. |
near_miss_false_positives |
near-miss cases where the skill fired. |
ablation_delta |
outcome pass rate with the skill minus without it, over the positive cases that have assertions or a rubric and ran in both arms. null without --ablation data. |
skill_tokens, run_tokens, cost_usd |
approximate tokens of SKILL.md; tokens and USD the runner reported. |
Rates are rounded to four decimals.
eval run records each skill in .ai-rulez/eval-results.json (commit it; it is deterministic, sorted by skill id,
and holds no prompts or outputs):
{
"schema_version": 1,
"skills": [
{
"id": "deploy-staging",
"digest": "sha256:...", // the skill's authored content when the run happened
"cases_digest": "sha256:...",
"lock_digest": "sha256:...", // the lock's digest of the skill; what usage logs join on
"cache_key": "sha256:...",
"runner": "claude-plugin-eval", "harness": "claude", "model": "haiku", "ablation": true,
"date": "2026-10-05", // from --date / $AI_RULEZ_EVAL_DATE, omitted when neither is set
"passing": true,
"score": { "cases": 3, "pass_rate": 1, "trigger_precision": 1, "trigger_recall": 1, "ablation_delta": 0.5, "...": "..." },
"last_pass": { "digest": "sha256:...", "date": "2026-10-05" }
}
]
}
The skill digest is the sha256 of every regular file under the skill directory except its top-level evals/, in path
order (path NUL length NUL bytes); editing a case does not make the skill look edited. It is the cache and
freshness key. lock_digest is the same skill under the lock's scheme (ai-rulez/skill/v1: SKILL.md
plus references/, scripts/ and assets/, again without evals/), the identity the usage log
carries as digest. A record without it (written by an earlier release) joins with usage by skill id only until the
skill is evaluated again; a cached run backfills it. last_pass survives a later
failing run, which is what freshness compares against.
Linting cases and results¶
These rules join AR962 in strict validation:
| Code | Name | Default | Reports |
|---|---|---|---|
AR996 |
eval-case-invalid |
error | A *.eval.yaml/*.eval.yml/*.eval.json file is malformed: unknown field, missing expect_trigger or prompt, bad assertion, invalid regex, unsafe path, duplicate id |
AR997 |
eval-stale |
off | The skill changed after its last recorded passing run. [lint.evals] require_fresh = "warn" or "error" turns it on |
AR998 |
eval-score-low |
off | The recorded pass rate is below [lint.evals] min_pass_rate (0-1); setting it turns the rule on at error |
AR9A0 |
eval-results-invalid |
error | eval-results.json cannot be parsed or has an unsupported schema_version |
AR9A1 |
activation-low |
off | The recorded activation recall or precision is below [lint.evals] min_activation_recall or min_activation_precision (0-1); setting either turns the rule on at error |
AR9A2 |
skill-confusable |
off | A sibling won at least [lint.evals] confusion_threshold (0-1) of the skill's positive activation prompts; setting it turns the rule on at warning |
AR9A3 |
activation-policy-conflict |
warning | A case expects a trigger (expect_trigger: true) for a skill whose frontmatter sets disable-model-invocation: true (or its disable_model_invocation misspelling) or allow_implicit_invocation: false: the model never starts it, so the case can never pass. A negative case for such a skill is reported too: it can never fail, so it measures nothing |
AR9A4 |
activation-prompt-names-skill |
off | A positive prompt contains the skill's id or name as a whole word (case-insensitive, /name included): it tests an explicit invocation, not whether the model chooses the skill. Enable with [lint.severity] AR9A4 = "warning" |
AR9A5 |
eval-import-unmapped |
info | Reported by eval import, never by validate: input fields with no counterpart |
Only records signed with your own key (eval-results.key in the user config directory, written by eval run) count
as evidence. A record without a valid signature (committed from another machine, edited by hand or forged) is
unverified: when AR997 or AR998 is enabled it is reported with an "unverified" message instead of being treated as
a passing run. Lint never creates the key, so on a machine that has never run eval run every record is unverified.
CI. A CI runner has no eval-results.key (it is per user and never committed), so every committed record is
unverified there. With require_fresh or min_pass_rate set, validate therefore reports AR997 and
AR998 as "unverified" for every skill that has a record, whatever the stored pass rate, and they fail the build at
the configured severity. Either run eval run in CI (it creates a key on the runner and records fresh, verified
results), or leave AR997 and AR998 off in CI (--lint-profile, or [lint.severity] AR997 = "off" in the CI
overlay) and enforce evals with the job that runs them.
[lint.evals]
require = true # AR962: skills need cases
require_fresh = "error" # AR997: edited after the last passing eval
min_pass_rate = 0.8 # AR998, and the default pass mark of eval run
min_activation_recall = 0.8 # AR9A1
min_activation_precision = 0.9 # AR9A1
confusion_threshold = 0.25 # AR9A2
[lint.evals.estimate] # assumptions of the cost estimate, see "Cost controls"
overhead_tokens = 25000
A skill with no recorded passing run is not reported stale (that is what AR962 and the score are for).
Reports¶
ai-rulez telemetry report evals joins the results with the usage log and the feedback log (see
Usage telemetry) and recommends an action per skill. The rules are fixed:
| Action | When |
|---|---|
rewrite |
pass rate below --min-pass-rate (0.8), trigger precision or recall below --min-trigger (0.8), a negative ablation delta, edited since the last passing eval, or misled+wrong+stale feedback outweighing great |
prune |
not rewrite; a usage log was given; never used in it; and evals do not show it helping (no record, or ablation delta of 5 points or less) |
review |
not rewrite or prune, but unused while evals show value, or no eval results yet |
keep |
everything else |
Rows are ordered by action, then by number of reasons, then by pass rate (worst first), then by SKILL.md size.
Records that are not signed with your key are marked unverified ("unverified": true in JSON), their scores are left out and the skill counts as having no eval results; telemetry report likewise skips them. Run eval run to record a signed result.
Without a usage log nothing is concluded about use. --format json prints the same data. telemetry report shows feedback
counts and the eval pass rate next to each skill.
With a usage log each row also carries a join class (and join_uses, the uses behind it by class) saying how far
the usage evidence is tied to the evaluated skill:
| Class | Meaning |
|---|---|
exact |
a use was logged at the skill digest the eval ran on (the record's lock_digest) |
stale |
the uses carry a digest, but none is the evaluated one: the score describes another version of the skill |
legacy |
the record has no lock_digest, or the uses have no canonical digest (version 2 log lines, an index without digests): matched by skill id only |
none |
the skill has no logged use, or no eval record |
The class does not change the recommended action; it says how much weight a row's usage deserves.
Bundling cases into a plugin¶
Eval cases are authoring material, so a plugin bundle leaves them out by default. (Before this option existed a
skill's evals/ directory was copied along with the rest of the skill directory; it is now excluded unless you opt
in.) Opt in with include_evals:
With it on, every runtime bundle that carries skills contains:
skills/<name>/evals/...for each bundled skill, andevals/..., a copy of.ai-rulez/evals/(or<content_root>/evals/whencontent_rootis set), at the bundle root next toskills/.
For a per-domain plugin ([marketplace.from_domains]) only evals/<skill-name>/... of the skills in that plugin is
bundled, so one plugin never carries another plugin's cases. The copy obeys the same rules as other bundled files:
regular files only, and a symlink that resolves outside the project is refused.
The provenance sidecar .ai-rulez-generated.json records a hash of every bundled file, cases included, so
ai-rulez verify --plugin fails when a case is edited or removed in the bundle.
Linting for skills without cases¶
evals-missing (AR962) is part of strict validation and is off by default:
[lint.evals]
require = true # turns AR962 on at warning severity
allow = ["internal-*", "scratch"] # skill names or globs exempt from the check
[lint.severity]
evals-missing = "error" # or set the severity directly; this wins over require
A skill passes when skills/<name>/evals/ or .ai-rulez/evals/<name>/ contains at least one non-hidden file.
.gitkeep does not count.
$ ai-rulez validate
.ai-rulez/skills/deploy-staging/SKILL.md:2 warning AR962 evals-missing skill "deploy-staging" has no eval cases ...
Running evals in CI¶
set -euo pipefail
ai-rulez validate # AR962/AR996/AR997/AR998 included when enabled; exit 2 on findings at or above fail_on
ai-rulez generate --plugin # writes the bundle, cases included with include_evals = true
ai-rulez verify --plugin # exit non-zero if any bundled file differs from its recorded hash
# Only the skills this change touched, with a JUnit report CI can display.
ai-rulez eval run --changed-only --base origin/main \
--ablation --max-cost 5 --date "$(date -u +%F)" \
--format junit --output-dir eval-report --runner-arg --trust-plugin
- Against the git diff.
--changed-only --base origin/mainruns the skills that have a changed or untracked file below their directory or below.ai-rulez/evals/<skill>/. Make the base ref available (git fetch origin main, or a full-depth checkout). Without--changed-onlyevery skill with cases is selected. - JUnit.
--format junit --output-dir eval-reportwriteseval-report/eval-report.xml: one suite per skill, one test case per eval case (near misses included), failures for failed cases, errors for runner errors and invalid case files, skipped for skipped or dry-run entries, and the scores as suite properties. The file has no timestamps or timings, so it is byte-stable. Point your CI's test reporter at it. - Caching by digest. Commit
.ai-rulez/eval-results.json, or restore it from your CI cache. A skill whose digest and cases digest match the stored run is reported ascachedwith its stored score and is not run again, so an unchanged skill costs nothing even without--changed-only(only records signed with your own key replay; see Caching). Commit the updated file (or save it back to the cache) after a run that changed it;validatewithrequire_freshthen fails the next change that edits a skill without re-running its evals. - Cost controls. Run
eval run --dry-runfirst to see the estimate; set--max-costso a runaway suite stops (estimate above the cap: refuse to start; spend reaching it: skip the rest, exit 2); keep--runslow in CI and--ablationon only where you track the delta; use a cheap--modelor per-casemodelfor smoke cases; a--dateis required for a dated result and comes from CI, never from the clock inside ai-rulez. - Exit codes.
eval runexits0clean,2for a failing/errored/over-budget/invalid skill,1when it could not run at all.validateexits0clean,1for an invalid configuration and2for findings at or abovefail_on;verify --pluginexits non-zero on any mismatch. - Secrets and trust. The runner runs the harness as the CI user with whatever credentials that job has. Evaluate
only skills and cases you trust;
command_exitassertions and--runner-arg --trust-pluginare opt-ins for that reason. ai-rulez itself makes no network call; the runner you choose does.