Improve (experimental)¶
ai-rulez improve run <skill> --with CMD runs an external optimizer (a tool that rewrites a skill against a
benchmark) on a throwaway copy of the skill and accepts a candidate only if a held-out eval set improves. ai-rulez
implements no optimizer. It owns the parts that make one safe to use on real sources: a disposable workspace,
a withheld held-out set, a strict acceptance gate, a diff policy, a cost ceiling and a human-reviewed apply step
or pull request.
Experimental. Flags, the optimizer protocol and
report.jsonmay change before this is stable. The command prints a warning on every run.
Quick start¶
# 1. tag at least three eval cases `holdout` (one negative), keep others as training cases
# 2. see the plan and the estimate; nothing runs
ai-rulez improve run deploy --with 'python optimize.py' --max-cost 5 --dry-run
# 3. run it
ai-rulez improve run deploy --with 'python optimize.py' --max-cost 5 --yes
# 4. review the run, then write it into the working tree...
ai-rulez improve show <run-id>
ai-rulez improve apply <run-id>
ai-rulez lock && ai-rulez eval run deploy && ai-rulez validate
# ...or open a pull request from an isolated worktree
ai-rulez improve pr <run-id>
The optimizer runs in the run's throwaway workspace, so relative paths in --with are made absolute from the
directory you run improve in: the program when it contains a path separator (./optimize.sh), and any later
argument that names an existing file there (optimize.py).
Commands: run, apply, show, clean, pr, adapters. Exit status of run: 0 candidate accepted, 2 no acceptable
candidate (report written), 1 refused or could not run.
Preconditions¶
A run is refused, with the reason, unless:
- The skill is authored (under
skills/or a domain), not from an include, an installed skill or a built-in. - Eval cases exist (
AR962) and parse (AR996). - At least three scored held-out cases exist, one of them negative (
expect_trigger: falseor a near miss) (AR9J2). - The skill and its evals have no uncommitted changes (checked when the project is a git repository; otherwise a warning).
--max-costis set and the eval estimate does not exceed it.- The secret scan is clean for the skill files and the train cases (prompts, fixtures, assertions,
rubricandrubric_items) the optimizer will receive.
Held-out split¶
Cases tagged --holdout-tag (default holdout) are held out. With no tagged case, hash(case id) mod 100 <
--holdout-fraction * 100 decides, deterministically. Near-miss expansions stay with their parent. A train prompt
that repeats a held-out prompt is reported as a warning.
The optimizer receives train cases only. Held-out prompts, assertions and rubrics are never written to the
workspace, the request, the environment or the logs. The optimizer learns only a one-word decision per round
(accepted, rejected: below gain, rejected: regression, ...), never per-case held-out results. Tests search
the request, environment, run directory and stderr for canary strings from held-out cases.
Optimizer protocol v1¶
The command runs without a shell, with the working directory .ai-rulez/local/improve/<run-id>/workspace/,
AI_RULEZ_IMPROVE_PROTOCOL=1, the request on standard input and the response on standard output (at most 1 MiB).
It edits the files in <skill>/ under the working directory.
{ "version": 1, "run_id": "imp-9f3a1c2d", "round": 1,
"skill": { "id": "deploy", "dir": "deploy", "digest": "sha256:..." },
"train_cases": [ { "id": "deploy-basic", "prompt": "...", "expect_trigger": true, "assertions": [] } ],
"train_scores": { "pass_rate": 0.67, "trigger_precision": 1.0, "trigger_recall": 0.5 },
"constraints": { "editable": ["SKILL.md", "references/**"],
"frontmatter_immutable": ["name", "allowed-tools", "disable-model-invocation", "model"],
"max_skill_tokens": 2400, "forbid": ["scripts/**", "assets/**"] },
"budget": { "max_cost_usd": 2.0, "timeout_s": 1200 },
"history": ["rejected: below gain"],
"previous": { "round": 1, "decision": "rejected: policy", "reasons": ["AR9J3 outside-editable: ... (scripts/run.sh)"],
"summary": "Added a helper script.", "workspace_kept": false } }
{ "version": 1, "summary": "Tightened the description.", "changed": ["SKILL.md"], "cost_usd": 0.42, "notes": "optional" }
previous (absent in round 1) lets a stateless optimizer avoid repeating itself: how the last round ended, the
optimizer's own summary of it, and whether the workspace still holds that attempt (a round the gate rejected) or was
reset to the state before it (a round rejected earlier, by the diff policy, the sibling guard or a failure). reasons
are given only for rounds decided before the held-out set was consulted; a held-out gate decision carries the one-word
decision and nothing else.
changed is informational: ai-rulez computes the real diff. cost_usd is self-reported and counts against
--max-cost; a run whose optimizer never reports a cost is flagged. summary and notes are untrusted: control
characters are removed and they are never fed to a model. A crash, timeout, invalid JSON or wrong version rejects
that round only, and the round is charged: the cost_usd of a well-formed answer the optimizer printed before it
failed (the bundled adapters print one), else the whole budget.max_cost_usd it was handed, so the next round gets
only what is left.
Bundled adapters¶
ai-rulez improve adapters lists them; improve adapters <name> prints a template. Select a runnable one with
--with builtin:<name> or --adapter <name>. An adapter is a child process of improve (ai-rulez improve adapter
<name>, hidden), so it gets the same sandbox, scrubbed environment, diff policy and held-out gate as any other
optimizer; nothing in the gate trusts it.
| Adapter | Kind | What it does |
|---|---|---|
builtin:noop |
runnable | changes nothing and reports no cost; for trying the protocol |
builtin:review-fix |
runnable | judges SKILL.md with the review rubric and applies a verified fix |
shell |
template | a shell script that calls any tool on the workspace copy |
research |
recipe | how to wire a research optimizer; its interface is unverified, so it is not shipped as supported |
repair-workflow |
template | a scheduled GitHub Actions workflow for model-upgrade repair |
builtin:review-fix drives the review pipeline of review fix on SKILL.md: the judge
(rubric builtin:skill-quality, content full) rates the skill; for stable findings the fixer proposes exact-text
edits that must pass the review checks (only the description and body change, growth at most [review.fix]
max_growth_percent and within the run's max_skill_tokens, no new link or credential, no new security finding, and a
judge of another model rates the targeted dimensions better and no other worse). A verified fix is written to the workspace; the gate then decides
whether it is kept. It never sees held-out cases (it uses none).
- Needs
[llm] modelandallow_network = truein the user config, the same opt-in asreview --semantic, and a fixer model different from the judge (--adapter-modelor[review.fix] model;--allow-same-modeloverrides, with the self-preference risk). Any of these missing refuses the run before anything starts (AR9J9). - The parent resolves the settings (user config, repository trust,
allowed_hosts) and hands the child the resolved[llm]table as an argument. A key is never passed: theapi_key_envvariable name is added to--env-pass, and the provider host (base_urlhost, elseprovider-default) is declared as egress unless you pass--egress. - Spend: the adapter's own model cost is reported as the optimizer cost, bounded by the run's remaining budget.
- A model that must be told apart from the judge cannot be the same alias: pin model ids.
- The adapter is a new process every round, so the request's
previousfield is its only memory. It hands the fixer, from the first attempt, how the last round ended (the decision, the reasons of a round decided before the held-out set was consulted, the adapter's own summary, and whether the file still holds that attempt), so a round after a rejection does not repeat it. This is theFixInput.Feedbackhook of the review fix API.
Diff policy¶
Checked after every round, before any eval spend. A violation is AR9J3 and rejects the round (the workspace is
restored to the last good state):
- a change outside
editable(defaultSKILL.md,references/**), including new, deleted or forbidden files; - a symlink, hard link, special or oversized file; an executable-bit change;
- a change to
name,allowed-tools,disable-model-invocationormodel(--allow-frontmatterfrees all butname); - a new reference to
scripts/(--allow-scriptsfreesscripts/**andassets/**); SKILL.mdabove[improve] max_skill_growth(default 1.25) times the original token count;- a new security finding (
AR0xx) in a changed file; findings already in the original never block; - a write outside the skill directory: the workspace root, the read-only original, the train cases or the authored skill.
A held-out assertion string that appears in the candidate but not in the original is reported as a warning (possible leak).
Sibling trigger guard¶
A rewritten description can pull prompts away from another skill. When a candidate changes what the find_skill
ranker reads (name, description, triggers, keywords), the trigger cases of every other skill of the project
are re-run with the offline ranker, once with the original and once with the candidate in the target's place. A
sibling whose trigger recall drops rejects the round as a regression (AR9J4), before any held-out spend and
without consuming a held-out evaluation. The report lists each sibling's recall and the prompts it lost.
The guard uses the retrieval surface of activation mode: no model, no cost, the same result everywhere. It is on whenever another skill has trigger cases (design question 5: mandatory, because it is free). It never copies the target skill's eval cases into its scratch tree.
A sibling that cannot be copied safely (a symlinked or oversized SKILL.md, too many eval files) is left out of both
arms and listed under unmeasured in the report, with a warning on the round; the guard measures the others instead
of failing every round.
Native surface (--sibling-native). The offline ranker is not the model that picks skills in a harness, so a
description that the ranker leaves alone can still pull a sibling's prompts away from the real model. --sibling-native
adds a second guard after the free one passes: the siblings' trigger prompts run through the harness's model
(claude-native, or a --runner-command that declares the activation capability and the native surface), each prompt
--sibling-runs times (default 3), once for the original, measured on the first round and reused, and once per
candidate. A sibling whose trigger recall drops rejects the round (AR9J4) exactly like the free guard, and the
result is siblings_native in the report. It costs money: every call is a model call, the spend (cost_usd of the
guard) counts against --max-cost, the run stops when no budget is left, and a guard that cannot run rejects the
round rather than clearing it. A body-only edit skips it. Model runs are noisy: raise --sibling-runs for a
large suite, and read a single regression with that in mind. The consent summary prints the guard when it is on.
Acceptance gate¶
Baseline and candidate are both measured in the run on the held-out set, --runs times each, with the majority
outcome per case. A candidate is accepted only if all hold:
- at least one held-out case was scored by both arms, and the candidate was scored on every case the baseline was
(a case the candidate could not be scored on counts as a loss, never as a pass; a gain of 0 with
--min-gain 0is not evidence); - held-out pass rate gain >=
--min-gainand at least one stable win (so--min-gain 0never accepts a candidate that gained nothing; unstable cases are not wins); - held-out cases flipping pass to fail <=
--max-regressions; - trigger precision and recall do not drop, and near-miss false positives do not rise;
- no sibling lost trigger recall (above);
- the diff policy held.
Cases that are not unanimous across runs are unstable: shown, and never counted as wins; a majority pass-to-fail
flip still counts as a loss, so noise cannot hide a regression from --max-regressions. A baseline that
already passes everything exits 2 before the optimizer runs. --max-holdout-evals bounds how often the held-out
set is consumed. Each round continues from the optimizer's own state; the best accepted round (highest gain,
earliest on ties) is kept, or the first with --stop-at-first-accept.
Confidence interval. Every comparison carries a percentile bootstrap 95% interval of the gain: 10,000 resamples
of the paired held-out cases, seeded from the case ids and outcomes so the same run always reports the same
interval. It is informative, not a gate: with 5 to 20 cases almost nothing is significant, and pretending otherwise
is worse than saying so. A run with fewer than 8 held-out cases, or an interval that includes zero, is marked
underpowered (AR9J5); the round warns about it only when it was accepted. Unstable cases count as no change in the
interval, as they do in the gate. --require-ci-above-zero (or [improve] require_ci_above_zero) makes the lower end of the
interval a gate condition, for suites large enough to clear it; the decision is then rejected: below confidence.
Cost¶
--max-cost is required and covers measured eval cost plus optimizer-reported cost. The pre-run estimate (shown by
--dry-run) must not exceed it. During the run the loop stops when spend reaches it (stopped: over budget) and
keeps the best accepted candidate. An eval runner that reports no cost is charged the whole remaining budget, as
eval run does. A candidate whose held-out measurement the budget cut short (fewer runs than --runs) is rejected
(rejected: over budget), and every cut-short measurement is named in the report's warnings.
Security¶
The optimizer is untrusted code running as you. It starts without a shell, in a scrubbed environment (PATH, locale, TZ
and names from --env-pass; HOME and the temp directories point into the run directory), with a timeout that
kills the process tree and capped output. A credential-like or proxy name in --env-pass is refused unless
--egress HOST declares where data goes; the declaration is printed in the consent summary and recorded in the
report, and is not enforced by ai-rulez itself.
Optional sandbox. --isolation auto|require (or [improve] isolation) wraps the optimizer in the process sandbox
of internal/sandbox: macOS sandbox-exec, Linux bubblewrap (or unshare, network only). Writes are limited to the
run's workspace, home and tmp directories, and the network is denied unless --egress declares a host. require
refuses to run when no backend can confine the process (AR9J7); auto confines when it can and says so when it cannot;
the default none keeps the optimizer unconfined, so existing optimizers keep working. The report records the mode, the
backend and what the backend enforces, and the consent summary states the same before anything runs: unshare (Linux
without bubblewrap) confines the network only, and --egress lifts the network denial. Reads are not restricted on any
backend: the sandbox does not hide the repository, the held-out cases or the per-user report key. improve pr takes
the same --isolation (and [improve] isolation) for the commands it runs in the worktree: see Pull request.
The held-out gate guards against an honest-but-overfitting optimizer, not a hostile one. The held-out cases are
files in the authored skill's evals/ directory, readable by absolute path by a process running as you, and the
optimizer is such a process. A hostile optimizer can read them, forge a run, or read the per-user key described
below. The key is a per-user file outside the repository, but it is readable by that same process, and improve apply
and improve pr re-check the diff policy against the live skill, not the held-out gate: a run forged by an optimizer
that read the key looks signed. Isolation none does not stop this; a sandbox does not either, because reads stay open.
Run improve run in a container (or CI job) with no access to the repository checkout beyond the run
workspace and an egress policy if the optimizer is not code you trust.
A repository must not choose what runs. [improve] optimizer and env_pass in a repository config (including
config.local.toml, which lives in the checkout) are used only with --trust-repo-optimizer; otherwise they are
reported and ignored (AR9J6). The same keys in the user config file are used without it.
A repository may tighten the gate, not weaken it. A repository [improve] table is honoured for min_gain,
max_regressions, holdout_fraction, max_skill_growth, runs, max_rounds and max_holdout_evals only when the
value is at least as strict as the default (gain >= 0.05, 0 regressions, held-out share >= 0.3, growth <= 1.25,
runs >= 3, rounds <= 3, held-out evaluations <= 3); require_ci_above_zero = true and a higher
min_holdout_cases always apply. A looser value is dropped with an AR9J6 warning unless --trust-repo-optimizer is
given. validate reports the repository keys improve run will not honour (AR9J6, a warning). The user config file may set any value, and so may a flag. The consent summary prints the effective gate:
minimum gain, regressions, held-out floor and share, growth limit and the interval requirement.
What leaves the machine: the skill text and the train cases, to whatever the optimizer calls. The consent summary
says so and needs --yes or an interactive confirmation. Held-out cases never leave ai-rulez.
[improve] configuration¶
Defaults of improve run; a flag always wins, the user config wins over the repository config.
[improve]
holdout_tag = "holdout"
holdout_fraction = 0.3
min_holdout_cases = 3 # never below 3; the split and the baseline need this many scored held-out cases
min_gain = 0.05
max_regressions = 0
runs = 3
max_rounds = 3
max_holdout_evals = 3
max_skill_growth = 1.25 # at most 2
require_ci_above_zero = false
isolation = "none" # none | auto | require
env_pass = [] # repository value used only with --trust-repo-optimizer
# optimizer = "..." # a command, or builtin:<adapter>; same trust rule
Output, show and clean¶
A run lives under .ai-rulez/local/improve/<run-id>/ (machine-local, git-ignored): plan.json, original/,
workspace/, train/, rounds/<n>/candidate/, report.json (improve-report/1) and, when accepted,
diff.patch (applies with git apply). report.json records the split by case id, every round's decision and
reasons, before/after scores with the bootstrap interval, the sibling guard, wins and losses, costs, declared egress,
the environment variable names (not values), the isolation, the adapter and the description change.
A run that stops on an error after measuring (a crashing eval runner, say) still writes its signed report.json with the
rounds and the spend so far and the reason stopped: ..., then exits 1.
improve show <run-id> prints the report and the diff (--format json prints improve-show/1). The optimizer's
text and the candidate's text (violation paths and details, refusal messages, round reasons, file names) are printed with control
characters and bidi controls replaced, so a candidate cannot write terminal escapes, and a candidate file whose name holds
such characters breaks the diff policy; apply does the same for its diff. improve clean <run-id> or --all deletes runs
(--dry-run; --all asks unless --yes); a symlinked run or improve directory is unlinked or refused, never followed.
Before each optimizer call the run directory outside workspace/, home/ and tmp/ (plan, original copy, train
cases, earlier rounds) is snapshotted. If the optimizer changed any of it, or the authored skill, the round is
rejected (outside-workspace) and everything is restored from the snapshot. A run id is imp-<8 hex> with an
optional -<n> suffix.
Apply¶
improve apply <run-id> shows the diff, asks for confirmation (--yes skips it) and writes the files. It refuses with
AR9J1 if the skill changed since the run (so a second apply fails), if the saved candidate no longer matches its
digest, or if the candidate now breaks the diff policy. It never commits. With --format json, stdout holds the
result document and the diff and prompts go to stderr.
improve run signs report.json with an HMAC under the per-user key (eval-results.key in the user config
directory, outside the repository, shared with eval run) and stores it as report.mac. apply and pr refuse a run
whose report is missing the MAC, was edited, or was made on another machine or by another user. The diff policy is
recomputed from the defaults and the flags, never taken from the report: edits to scripts/ and assets/ need
--allow-scripts, and changes to allowed-tools, model or disable-model-invocation need --allow-frontmatter,
even if the run itself used them. Files are written all or nothing.
Pull request¶
improve pr <run-id> [--base BRANCH] makes the change reviewable without touching your checkout:
- Creates a linked git worktree (in a temporary directory) on a new branch
ai-rulez/improve/<skill>-<digest8>started from--base(default: the branch you are on). A detached HEAD needs--base. An existing branch of that name is refused. - Checks that the skill at the base has the digest the run measured (
AR9J1otherwise), that no component of the path to it is a symlink, re-checks the diff policy (with the growth factor the run used) and writes the candidate there. - Runs
ai-rulez generate --yesin the worktree, so the outputs the lock pins are not stale; when git tracks any file of the generate manifest (the project commits its outputs), all of them are staged with the skill. A project that ignores its outputs gets none. A failing generate stops the pull request when the project has a lock and only warns without one. Then it runsai-rulez lockwhen the project has a lock. These commands run with a scrubbed environment (the base variables andXDG_*/AI_RULEZ_HOME; credentials only by name with--env-pass, which an eval runner needs). With--run-evalsit also runsai-rulez eval run <skill> --changed-only(plus any--eval-arg) soAR997holds; that calls the eval runner and spends money, so it is off by default and the command says what to run on the branch instead. A failing step stops the pull request and leaves no branch. - Stages the skill, the lock and the eval results, and makes one commit (
chore(skills): improve <skill> ...; git hooks are skipped). The worktree is removed; the branch stays. - When the remote exists,
ghis onPATH, the base is a branch the remote has with no unpushed commits (a commit id, an unpushed branch or a base ahead of the remote would break or pollute the pull request), and you confirm (or pass--yes), pushes withgit push --set-upstreamand runsgh pr create [--repo OWNER/REPO] --base B --head BRANCH --title T --body-file F [--draft];--reponames the repository the remote URL points at, so a fork receives the pull request on the fork. Otherwise (no remote, nogh,--no-push, not confirmed, an unusable base) it prints the two commands. The body file is written before the worktree exists, so a write failure leaves no branch. improve pr itself makes no network call, but thegenerateandlockit runs in the worktree fetch remote includes and sources as they do anywhere (andeval runcalls the eval runner); git andghuse their own credentials, andghgets a scrubbed environment plus the usualGH_*/proxy names.
Isolation. improve pr --isolation auto|require (or [improve] isolation) runs generate, lock and eval run in the
worktree under the same process sandbox: writable only below the worktree's project directory, ai-rulez's own
subdirectories of the user cache (~/.cache/ai-rulez, $XDG_CACHE_HOME/ai-rulez) and data ($XDG_DATA_HOME/ai-rulez)
roots, AI_RULEZ_HOME and the temp directory, where they exist (the two subdirectories are created). The network is
cut for generate and lock; --allow-network leaves it on for a project whose remote includes are not cached yet, and
eval run (--run-evals) keeps it because it calls a model. The candidate is the
optimizer's text, and eval run lets an agent act on it, so a hostile candidate cannot write outside those
directories. require refuses (AR9J7) before the worktree is made when no backend works; auto warns and runs
unconfined. The result (isolation in --format json) and the first line printed name the backend and what it enforces.
A harness that keeps state elsewhere (~/.claude, say) cannot run --run-evals confined; use --isolation none for that step.
What the pull request step runs, and with what. Nothing goes through a shell: git and gh start with argument vectors,
branch, remote and path arguments are checked for option injection, and commits skip hooks (--no-verify).
Every git command also runs with core.hooksPath pointing at the null device and core.fsmonitor=false, so no
hook of the repository runs for the worktree, add, commit or push. Git gets a scrubbed environment (the base set, its
own configuration and the commit identity variables, never a model key or GH_TOKEN), and only the push adds the
transport credentials it needs (SSH_AUTH_SOCK, GIT_ASKPASS, SSH_ASKPASS, GIT_SSH, GIT_SSH_COMMAND; an HTTPS
credential helper is read from git's own configuration). gh gets its own scrubbed environment, and the
ai-rulez children get the --env-pass names only. Every process is bounded: 30 minutes for each ai-rulez step, 2 minutes
for gh, and the runner's ceiling for git. The push is the only step that reaches the network without being asked.
The body states what changed, the held-out table with the interval and the wins and losses, the guards that held, cost,
egress, the environment names forwarded, the description before and after, the optimizer's own words (untrusted: inside
code fences or code spans longer than anything they contain), a reviewer checklist, and that the change is not
approved. Nothing here sets approval. Refusals carry AR9J8.
Scheduled repair workflow¶
A model upgrade can make a skill that passed its evals fail. ai-rulez improve adapters repair-workflow prints a
GitHub Actions workflow (save it as .github/workflows/ai-rulez-repair.yml) that does the loop on a schedule:
eval run --model $EVAL_MODEL --max-cost $EVAL_BUDGET --format json --no-writefinds the skills that now fail (jqoverskills[]withpassing: false); exit status 1 (could not run) fails the job.- For each one,
improve run <skill> --adapter review-fix --sibling-native --max-cost $REPAIR_BUDGET --yesrepairs it under the usual gate: held-out set, sibling guards, diff policy, cost ceiling. - A run that exits 0 gets
improve pr <run-id> --draft --yes: a draft pull request from an isolated worktree. A run with no acceptable candidate (exit 2) opens nothing.
Nothing is merged, approved or applied for you, and the pull request body says the change is not approved. The template
has no pull_request trigger, so a fork's pull request never reaches its secrets; its permissions are contents: write
and pull-requests: write for the branch and the pull request. It spends money in three places (the eval pass, each
repair, the optimizer's model), each with its own ceiling in the env block; the builtin:review-fix step also needs
the [llm] user config the template writes (delete that step and use --with for an optimizer of your own). The run
state lives in .ai-rulez/local/ of the runner, and improve pr needs the report MAC of the machine that made the
run, so both commands must run in the same job, as they do here. A test keeps every flag the template passes in sync
with the CLI.
Codes¶
| Code | Name | Meaning |
|---|---|---|
AR9J1 |
improve-run-stale |
the skill changed since the run (or differs at the base of pr) |
AR9J2 |
improve-no-holdout |
too few held-out cases |
AR9J3 |
improve-policy-violation |
a candidate round broke the diff policy |
AR9J4 |
improve-sibling-regression |
a candidate lowered another skill's trigger recall |
AR9J5 |
improve-underpowered |
the held-out gain cannot be told from zero |
AR9J6 |
improve-repo-optimizer-ignored |
a repository [improve] optimizer or env_pass was not used |
AR9J7 |
improve-isolation-unavailable |
the requested isolation could not be applied |
AR9J8 |
improve-pr-refused |
improve pr refused |
AR9J9 |
improve-adapter-refused |
a bundled adapter could not run |
They appear in improve output and the report only; validate never emits them. See strict validation.
Design decisions¶
- The gate is the product. The optimizer wins only by improving a score it never saw, without regressions, within budget.
- Gate on paired no-regression, report intervals. Suites of 5-20 cases cannot clear a significance test, so teams would
switch the gate off. The interval is reported and flagged;
--require-ci-above-zerois opt-in. - Baseline re-measured every run, never read from
eval-results.json, to avoid comparing across model versions. - No auto-merge, no auto-commit. The gate reduces review effort; it does not replace it.
improve prcommits only on a new branch of a throwaway worktree and states the change is not approved. - Run state is machine-local. Nothing durable is written to the project until
applyor a pull request. - Subcommand
run.improve run <skill>keepsimprovea pure command group; a skill namedapplycannot collide withimprove apply. - Held-out marking is the
holdouttag (design question 1), which works today; asplit:field can replace it later. --env-passhas a config equivalent only from trusted scope (question 2):[improve] env_passin the user config, or in a repository config with--trust-repo-optimizer.improve prdepends on theghCLI (question 3), not on a GitHub client of ai-rulez: less code and no extra egress to audit.- No research optimizer is documented as supported (question 4) until someone verifies its interface; the
researchrecipe says so. - The sibling guard is mandatory when a sibling has trigger cases (question 5): it is free, because it uses the offline ranker.
- Self-generated held-out cases (question 6) are not detected; the report cannot tell who wrote a case. Treat a model-written held-out set with the scepticism the design notes call for.
- The sandbox is opt-in (
--isolation, onrunand onpr): turning it on by default would break optimizers that write caches elsewhere or reach a local model server without declaring egress, and harnesses that keep state outside the worktree. - The native sibling guard is opt-in (
--sibling-native): it needs a runner with the activation capability and costs money, so the free offline guard stays the default. - Not done: the native guard for harnesses other than
claudewithout a--runner-command.
The accepted candidate is the best of the rounds that were scored on the held-out set, so with more than one such round
the gain is optimistic. improve apply and the pull request body say "selected among N held-out evaluations" then;
max_holdout_evals bounds N.