Performance¶
How ai-rulez is benchmarked, how a regression is caught, and what the optimization passes changed.
Running the benchmarks¶
Every benchmark that reads a project builds a synthetic one in a temporary directory
(internal/testutil.BuildBenchTree): rules, context, skills, agents and commands with frontmatter and prose, plus
a git repository when the code under test asks git questions. Nothing touches the network.
| Size | Content files | Runs by default |
|---|---|---|
| small | 10 | yes |
| medium | 200 | yes |
| large | 2000 | with AI_RULEZ_BENCH_LARGE=1 |
task test:benchmark # every benchmark, small and medium
task test:benchmark:large # adds the 2000-file project, one iteration each
go test ./internal/lint -run '^$' -bench RunFull -benchmem -cpuprofile cpu.out
go tool pprof -top -cum cpu.out
| Package | Benchmarks |
|---|---|
internal/parser |
ParseFrontmatter (no/small/rich/CRLF/long body), ParseFrontmatterBatch |
internal/tokens |
Estimate, Count (cl100k and byte ratio, prose and long-line input) |
internal/config |
LoadConfig |
internal/includes |
LoadWithLocalInclude |
internal/lint |
LoadTree, RunFull, Analyzers (each analyzer alone), SecretScan (DetectSecret, RedactSecrets, ScanText), SecretScanLongLine, SecretScanCustomPatterns, RunSkillTree, AnalyzerSelection, NumberedPrefixExists |
internal/contentlock |
Compute, BuildAndCompare (lock digest, then verify) |
internal/lockfile |
SaveLoad |
internal/generator |
Generate (full and idempotent no-op), GenerateLLMsTxt, Clean (remove and dry run), CheckDrift, SecretGate (plain and -memo, the context generate runs under) |
internal/review |
Score (collect and score), Calibrate |
internal/mcp |
ProjectTools (tool dispatch), SkillServer (catalog build, search_skills, load_skill) |
internal/verifiers |
RegexFiles |
cmd/commands |
Lock (write and check), PublishDryRun (bundle, agent-plugins, ard), OKF (export, validate) |
Regression tracking¶
scripts/bench.sh runs a curated subset (benchmarks that start no process and do not fsync, so they measure the code
and not the machine) with -count=6 and compares the median time/op of each benchmark with
scripts/bench-baseline.txt. It exits 1 when any benchmark is more than 25% slower. When benchstat is installed
(or the machine is online, through go run golang.org/x/perf/cmd/benchstat@latest) its report is printed too.
scripts/bench.sh # run, compare, fail on a >25% regression
scripts/bench.sh --update # re-record the baseline on this machine
Time per operation depends on the CPU, so the gate only applies when the baseline was recorded on the same CPU model
as the run (the cpu: line of the output). Otherwise, or when there is no baseline, it prints a warning and exits 0.
The CI Benchmarks job runs the script after the full benchmark run and uploads bench-subset.txt; to make the gate
bite on the CI runners, commit that artifact as scripts/bench-baseline.txt.
Optimization passes¶
Numbers are time/op and B/op on an Apple M5 Pro that was running other builds at the same time (load average near 50), taken as alternating runs of test binaries built from the commits below, so the ratios are reliable and the absolute values are not. Behavior is unchanged: the golden tests and every existing test pass, and each change that could alter a result has a test comparing it with the code it replaced.
| Path (benchmark) | Change | Before | After |
|---|---|---|---|
Credential scan of clean text, 37 KB (lint SecretScan/DetectSecret/clean/large) |
each credential pattern runs only on text holding one of its literals | 6.1 ms | 0.30 ms (20x) |
Credential redaction of clean text (SecretScan/RedactSecrets/clean/large) |
same gate in secretpat.Redact and lint.RedactSecrets |
6.1 ms, 0.9 MB | 0.28 ms, 0 B |
Redaction of text holding secrets (SecretScan/RedactSecrets/dirty/large) |
same gate | 6.1 ms, 0.9 MB | 2.5 ms, 0.33 MB |
Security scan of one text (SecretScan/ScanText/clean/large) |
credential gate, then literal gates on the shell-exec (curl, wget, base64, eval, chmod, >) and instruction-override patterns |
29 ms | 7.6 ms (3.8x) |
| Same scan, allocations | ASCII-only lines skip the rune conversion and seen-set of the hidden-character scan | 1.33 MB | 1.18 MB (-11%) |
Configured secret patterns, 400 lines x 2 patterns (SecretScanCustomPatterns) |
compiled once per run, not once per scanned line | 12 ms, 7.7 MB | 8 ms, 1.1 MB |
Lint of 200 content files (lint RunFull/medium) |
all of the above, tree-relative paths memoized per run, backtick-token prose check only for path-like tokens | 570-660 ms, 53.5 MB, 667k allocs | 265-335 ms, 47.1 MB, 617k allocs |
Frontmatter parse of a 570 KB document (parser ParseFrontmatter/longbody) |
find the closing marker by scanning instead of splitting and re-joining the document | 83.6 us, 228 KB | 10.2 us, 7 KB (8x) |
Frontmatter parse of 200 documents (ParseFrontmatterBatch/medium) |
same | 1.29 ms, 2.09 MB | 1.0 ms, 1.59 MB |
Any-file verifier regex, 1000-file tree (verifiers RegexFiles/medium) |
exclude globs compiled once per predicate, not once per matching file | 13.2 ms, 19.8 MB, 174k allocs | 1.3 ms, 61 KB, 238 allocs (10x) |
The shell and instruction-override pre-checks compare ASCII literals without case in place and treat U+017F and
U+212A as the s and k the patterns' case folding makes them, so a line the pattern would match always reaches it.
Second pass: repeated loads, publish at scale, startup¶
Same method as above, taken on a quieter machine (load average near 20): test binaries built from the commit before
the pass and from the last commit of it, run back to back with -benchtime=3x -count=2 (1x for the 2000-file
project), so the ratios are reliable and the absolute values are not. Output is unchanged: tests/golden passes
without edits.
| Path (benchmark) | Change | Before | After |
|---|---|---|---|
lock --check, 200 files (commands Lock/check/medium) |
the configuration is loaded once and shared by the content comparison, the tag check and the signature check (it was loaded three times); one repository-question memo per check | 1.5 s, 71.4 MB, 274k allocs, 16 git processes | 0.45 s, 45.1 MB, 192k allocs, 10 git processes |
lock --check, 10 files (Lock/check/small) |
same | 29 ms, 2.19 MB | 9.6 ms, 1.36 MB |
lock --check, 2000 files (Lock/check/large) |
same | 596 MB, 2.24M allocs | 393 MB, 1.69M allocs |
publish --dry-run, 10 files (PublishDryRun/bundle/small) |
the load is shared by the lock check, trap files are read at their size instead of into a 1 MiB buffer, the source hash is streamed into the digest, and one memo serves the whole run | 1.49 s, 49.6 MB, 27 git processes | 0.66 s, 37.7 MB, 22 git processes |
publish --dry-run, 200 files (PublishDryRun/bundle/medium) |
same | 4.3 s, 280 MB, 770k allocs | 2.2 s, 152 MB, 656k allocs |
publish --dry-run, 2000 files (PublishDryRun/bundle/large) |
same | 2.51 GB, 10.5M allocs | 1.36 GB, 9.89M allocs (-46% bytes) |
Numbered-path reference check (lint NumberedPrefixExists, 100 lookups over 20,000 paths) |
the paths are sorted once per tree and each lookup is a binary search, instead of listing and scanning every path per numeric token | 181 ms, 33.2 MB | 0.04 ms, 4.9 KB |
Generator secret gate (generator SecretGate/small, medium) |
generate runs it under a context that remembers git rev-parse answers, so the gate asks git once, not twice |
3.5 MB, 13.4k allocs; 3.7 MB, 15.0k allocs; 2 git processes | 1.9 MB, 7.4k allocs; 2.1 MB, 8.6k allocs; 1 git process |
Process start (ai-rulez --version, GODEBUG=inittrace=1) |
package-level regular expressions of lint, evalimport, cmd/commands and secretpat compile on first use (sync.OnceValue) instead of at start |
11.7 MB allocated by package init (6.8 MB in ai-rulez packages) | 9.1 MB (4.2 MB) |
What changed where:
lock --checkandpublish --dry-runload once and thread the result. The memo (gitutil.WithMemo) is attached to the context of onegenerate,drift,lock --checkorpublishrun and dies with it; it is never process-wide, and the long-lived callers (the MCP server, the doctor, watch mode) do not get one. Questions about history (HEAD, refs, file lists) are not memoized.- The 2.5 GB of
publish --dry-runat 2000 files was mostly one allocation: the trap check read every candidate file into a fresh 1 MiB buffer. The source hash text (built whole, hashed, dropped) and the repeated configuration loads of the lock check were the next removable allocations. - Package-level regular expressions in the four packages are lazy. The
globalarchlint rule treats async.OnceValuethe way it treatsregexp.MustCompile(immutable in practice); the three existingOnceValueglobals (inard,tokensandgenerator/settings) lose one allowlisted site each.
Behavior that changed deliberately (everything else is byte-identical):
- Frontmatter is split by
internal/frontmatterin the config loader (parseFrontmatter,hasUnclosedFrontmatter,FrontmatterProblem,IsOKFListing), in lint (parseDoc) and in the generator (frontmatterEnd, hash injection,listingFrontmatter,ruleFileFrontmatter,mappedFrontmatterFields). Its rules now apply to all of them: a UTF-8 BOM before the opening fence is ignored (the loader and the generator used to treat the file as having no frontmatter); a fence may carry trailing spaces or tabs, and an indented---is not a fence (the loader and lint used to trim both sides); the empty block---/---is a block (the loader needed three lines); the blank line after a CRLF block is dropped from the body as it is after an LF block;ruleFileFrontmatterno longer skips blank lines before the opening fence and no longer takes---foofor a closing fence; hash injection into a block closed on its first line works.internal/frontmatter.ClosingLinefinds the closing line without building the block, so lint does not copy the YAML. - Readers now refuse input over their limit through
safefs.ReadLimitedinstead of passing a cut one on: an attestation file overMaxBundleBytes(AR721 at the read, not at the parse), the GitHub token and Sigstore service responses over their caps, and the publish inputs (templates, the lock, skill files, dist files and tar entries), the signing key, the LLM cache entry and secret.safefs.ReadFileLimitedopens and reads with the same bound.
Third pass: the process cleanup after each spawn¶
internal/runner ends every run by killing what the command left behind. The guarantee, from the code and the
tests (TestRunKillsHelpersThatLeaveTheGroup, TestRunKillsADetachedSystemBinary, TestConcurrentRunsDoNotKillEachOther):
after Run returns, no process of the run is alive. That covers the process group, any descendant seen while the
command ran (a helper that called setsid or setpgid is still a descendant), and a helper that detached and was
reparented to init between two looks at the table, which is found by the run token every child inherits in its
environment or by the write end of the marker pipe every child inherits as a descriptor (macOS withholds the
environment of platform binaries, so there the descriptor is what finds them). It does not cover a helper that drops
both the variable and the descriptor before detaching, and it must never touch a process of another run. On Windows
the job object gives the guarantee without any table read.
The cost was in the token and descriptor search (adoptCarriers): it skips processes older than the command's root
(a descendant cannot be older than its ancestor), but a Spec.ShortLived run (every git call) does not read the
process table at start, so it never learned the root's start time and the search covered every process of the user,
one sysctl (macOS) or /proc/<pid>/environ read (Linux) each, after every spawn.
Now a short-lived run reads the start time of its own root with a single lookup at attach (processStart: one
sysctl kern.proc.pid, or one /proc/<pid>/stat), so the search covers only processes that started since the command,
which is a handful. If that lookup fails the search is unbounded as before, so the guarantee is the same on every
platform; nothing was weakened. The full (non-short) path already had the bound and is unchanged. Windows is unchanged.
Measured on a loaded machine (load average 16 to 28, 850 processes), go test -bench of the runner and generator
packages, test binaries built before and after and run back to back. Wall times are dominated by the fork cost under
that load and are not reliable; the allocations and the isolated sweep are.
| Path (benchmark) | Before | After |
|---|---|---|
Sweep alone over the real table of 850 processes (runner AdoptCarriers) |
9.5 ms, 1.5 MB, 10k allocs | 0.02 to 0.04 ms, 544 B, 9 allocs |
One short-lived spawn with cleanup (runner SpawnCleanup/short) |
1.85 MB, 7.3k to 7.8k allocs | 0.78 MB, 300 to 420 allocs |
One full spawn with cleanup (SpawnCleanup/full) |
unchanged | unchanged |
Secret gate, 10 files (generator SecretGate/small) |
4.2 to 4.5 MB, 18k to 19k allocs | 2.1 MB, 3.7k allocs |
Secret gate with the memo (SecretGate/small-memo) |
2.4 to 2.5 MB, 10k to 11k allocs | 0.9 to 1.2 MB, 0.8k to 2.4k allocs |
The cost of the sweep now depends on how many processes start during the command, not on how many exist.
What dominates elsewhere¶
These were measured and not changed here.
- Process spawns per content file. Loading a project with 200 skills took 91 s on the loaded machine before the
branch that batches
git ls-filesandgit check-ignoreper directory (0.5 s with it). The remaining wall time ofgenerate,clean,lockandpublishon 200 to 2000 files is mostly git (rev-parse,ls-files) and file-system calls; their CPU time is a small share. - Token counting. cl100k counting of the prompt text is the largest allocator in a lint run (about a third of
the bytes, in
regexp2match objects); the memo that removes it is part of the same separate branch. - Served skills.
lock --checkstill builds the served-skill views twice (mcp.DynamicLockChangesloads the configuration for each view and scans the skills); that code lives ininternal/mcp. - Publish at scale. What remains of
publish --dry-runat 2000 files (1.36 GB over 9.9 million objects) is the strict lint run (about 60%, mostly cl100k token counting inregexp2, which the separate token memo branch removes), the lock comparison and the served views.