Catalog¶
ai-rulez catalog describes everything a repository defines: rules, context, skills, agents, commands and
checks, with owner, version, size, the digest ai-rulez.lock pins, the roles that keep each item and the lock
status. It prints a table, JSON for tools, or a static website.
JSON¶
ai-rulez catalog --format json # version 1 (default)
ai-rulez catalog --format json --schema-version 2 # version 2
Version 1 (schema/catalog.v1.schema.json)
stays the default for one minor release so existing consumers keep working. Version 2
(schema/catalog.schema.json) is a
strict superset of the version 1 fields, plus:
| Field | Meaning |
|---|---|
generated_by, project |
tool name and version; project name, description and lock tree digest |
items[].ref |
stable key kind/domain/id (domain - when none); a repeated key gets #2, #3 |
items[].description, source |
description from the frontmatter (verbatim; a consumer escapes it); local or include |
items[].load_cost |
listing_tokens (what the harness always lists: skill name and description), body_tokens (loaded on use), resource_tokens and resources (bundled files) |
items[].lint |
status (ok, warn, error), counts and findings, from the same engine as validate; absent when lint could not run |
items[].excerpt |
first 2 KiB of the body, plain text; off with --include-excerpt=false |
items[].approval |
null, or {required, status, reviewers, assurance, expires} when [governance] requires approval of the item or the lock records one; see Approvals |
items[].eval, items[].usage |
with --with-eval / --with-usage: a skill's recorded eval result and use count, see Eval and usage |
mcp_servers |
the project's MCP servers: ref, name, transport, command_basename, enabled, profiles, pinned (exact version or digest in a package-runner launch; null when not applicable), and the names of env and headers entries, each marked literal or with the variable it references (ref); warnings flags an unpinned launch or a credential written as a literal |
edges |
dependencies between items: {from, to, kind: "uses"}, where from names the skill to in its skills: frontmatter (both are item refs); roles are not edges, see items[].roles |
lint |
project totals, counts per code, and the findings no item owns |
notes |
why a section is missing or narrowed |
MCP servers never expose arguments, URLs, env values or header values: a catalog is published, and those carry launch secrets and internal hostnames. Only the executable's file name and the names of env and header entries appear.
Paths are relative to the configuration directory; no absolute path of the machine appears. A consumer must refuse
a schema_version it does not know. The MCP catalog tool prints version 1, equal to the CLI default.
Eval and usage¶
ai-rulez catalog --format json --schema-version 2 --with-eval --with-usage
ai-rulez catalog --html site/ --with-eval=path/to/eval-results.json --with-usage=path/to/usage.jsonl
Both are off by default. Without a value they read the project's own files, eval-results.json in the configuration
directory and local/usage.jsonl (the log the usage hooks write); --with-eval=FILE / --with-usage=FILE name
another file (a file called default needs ./default). A missing file is not an error: the catalog's notes
say eval-results.json not found: eval fields are omitted, and nothing is invented. A file that cannot be parsed is
an error.
Only an allowlist of aggregates is copied:
eval:cases(scored),pass_rate,passing,ablation_delta,trigger_precision,trigger_recall,stale(the skill's lock digest differs from the one the run recorded),verified(the record carries a valid signature of this machine's key; a result committed from another machine is shown but unverified) anddate.usage:invocationsandlast_seen(a day). Sessions, harnesses, outcomes, feedback and notes never enter the catalog.
The log names a skill by id only, so the use of two skills that share an id is left out rather than guessed (the notes say so). The overview table gains an Eval and a Uses column when any skill has the data; each item page gets an Eval and a Usage table.
Dependency graph¶
The site has a graph.html page: items that name a skill in their skills: frontmatter (agents, skills, rules)
drawn as an SVG, left to right, so an item sits to the left of the skills it uses (longest path layering; roles
are drawn in a first column with the items they keep). It is a hand-laid, deterministic drawing: integer
coordinates, no script, no external library, no url() references, and the same bytes for the same catalog. Every
node links to its item page. The same edges are listed in a table below the drawing, which is the accessible
alternative to the picture.
A name resolves to the skill of the same domain, else to the root skill, else to the only skill of that name. A name
that matches nothing, or several skills in other domains, is left out and said in notes: it is never guessed. A
dependency loop is drawn dashed and listed. A graph over 150 items is shown as the table only, and role edges are left
out past 300.
Comparing catalogs¶
ai-rulez catalog diff main # a revision against the current project
ai-rulez catalog diff v5.0.0 HEAD --format json
ai-rulez catalog diff before.json after.json --exit-code
Prints what was added, removed or changed between two catalogs: items (matched by ref; a change lists the fields
that differ: digest, description, owner, version, path, source, delivery, listing/body/resource tokens, lint status,
approval status, roles), MCP servers (transport, command, pin status, env and header names), roles, dependency edges
and the lint totals. JSON output validates against schema/catalog-diff.schema.json.
Each argument is a catalog JSON file (catalog --format json --schema-version 2, or a site's catalog.json;
a schema_version other than 2 is refused) or a git revision. A revision is read without touching the working
tree: git archive writes the tracked files of the configuration directory at that commit into a temporary
directory (bounded in file count and size, extracted through a root so nothing lands outside it, symlinks reported
and skipped) and a catalog is built from it. With one argument the other side is the current project. Both sides are
built from the shared configuration only: the machine-local overlay is left out, and remote includes and installed
skills are not resolved (no network), so compare two catalog.json files to cover them. Untracked and ignored files
are not part of a revision. Both sides are linted against the working tree's repository, so a path the content names
(AR401) is checked the same way on each side, and the plugin version drift check (AR961), which compares with
outputs generated on disk, is left out of both.
Exit code 0 unless the command could not run (1); --exit-code makes it 2 when the catalogs differ. Text output
escapes control and bidirectional characters from the (possibly third-party) JSON.
Static website¶
Writes a directory with an overview (search by name, domain, owner, kind and lint status), one page per item and
per role, the MCP servers, the dependency graph, the lock status, the lint findings, an About page, catalog.json (the version 2 document the pages are
rendered from), assets/catalog.css, assets/catalog.js and robots.txt.
- Offline. All links are relative, so the site works from
file://and under any URL path. Nothing is fetched: no fonts, CDN, analytics orfetch. Every page is complete without JavaScript; the script only adds filtering and copy buttons. - Reproducible. The same input gives the same bytes: no timestamps, no absolute paths, sorted everywhere. The
footer shows the tool version and the sha256 of
catalog.json, not a date. Regenerate and diff to detect a tampered hosted copy. - Escaped. Every string from the repository goes through
html/templatecontextual escaping. Links are built from slugged keys, never from source text; source URLs are text. Invisible and direction-changing characters (bidi controls, zero-width characters) are replaced with U+FFFD and the item is marked "contains hidden characters". Bodies are shown as plain text in<pre>, never rendered. - Content-Security-Policy. Each page carries
default-src 'none'; img-src 'self' data:; style-src 'self'; script-src 'self'; base-uri 'none'; form-action 'none'in a meta tag, and no inline script or style. A host should send the same as response headers plusframe-ancestors 'none', which a meta tag cannot set. - Output directory. It must be new, empty or contain the marker
.ai-rulez-catalog(written by an earlier run, listing each file it wrote with its SHA-256).--cleanremoves only a listed file that has the shape of a site file (index.html,items/,roles/,assets/,catalog.json, ...) and still holds the recorded bytes; without the marker the run is refused, so--html .cannot overwrite a project, and a directory with a.gitor.ai-rulezfolder is always refused. The marker is written before the files, so an interrupted run leaves the directory marked. Writes never follow a symlink out of the directory. - Secrets. The run is refused when the secret scanner (
AR001) flagged an item and the site would publish its excerpt or description; remove the secret, or pass--allow-findings AR001(discouraged). The published description and excerpt are also scanned directly, so a lint that did not run, a baselined finding or an inline ignore does not let a secret through.
Configuration and pages¶
[catalog]
title = "Acme catalog" # site title (--base-title)
include_excerpt = true # body excerpts (--include-excerpt); default on, off when indexable
exclude_owners = false # leave owner names out of the JSON and the site (--no-owners)
indexable = false # let search engines in (--indexable)
max_items_per_page = 200 # overview rows per page (--max-items-per-page); 0 means 200
render_markdown = false # render excerpts as sanitized Markdown (--render-markdown)
Every key is optional and a flag that is given wins. include_excerpt and exclude_owners also apply to
--format json --schema-version 2. The [catalog] table can be overridden in the machine-local config.
The overview lists max_items_per_page rows per page. All rows are in the one index.html (one <tbody> per page),
so browsing, find-in-page and the no-JavaScript view show everything; the script shows one page at a time with
Previous/Next buttons and shows every page while a filter is active. The page size does not change catalog.json.
Markdown excerpts¶
--render-markdown (or render_markdown = true) shows each item's excerpt as formatted Markdown instead of plain
text. The renderer is a sanitizer by construction, not by filtering: the text is parsed as CommonMark and turned into
a tree whose nodes are only paragraphs, headings (shifted to h3-h6, below the page's own headings), lists, quotes,
code blocks, emphasis, code spans, line breaks and rules. The tree has no field that holds markup, and the page is
built from it by html/template, so every character of the body is escaped.
- Raw HTML (block or inline, comments included) is shown as literal text, never interpreted.
- Links are not anchors: the label is kept and the destination is shown as text after it, so a
javascript:,data:or remote URL never becomes anhref. Images show their alt text and the destination as text; nothing is loaded. - Depth is capped at 12 levels and the tree at 4,000 nodes (a notice says so), so a hostile body cannot recurse or inflate the page.
- The Content-Security-Policy stays as strict as before; the Markdown view needs no script and no new source.
The excerpt is still the first 2 KiB of the body, so a long document is cut mid-structure. Excerpts are off with
--indexable unless asked for.
Freshness check¶
Renders the site in memory and compares it with the directory without writing anything. Exit 0 when every file
matches and nothing else is there, 2 when a file is changed, missing or unexpected (the differences are listed),
1 when the check could not run. A directory that does not exist is drift. Because the output is reproducible, the
same gate detects a tampered hosted copy. Use it in CI to keep a committed site current; --check and --clean do
not combine. Files are read through the directory only: a symlink in it is reported, never followed.
Flags: --role R (items role R keeps), --include-excerpt (default on), --indexable (no robots.txt, no
noindex; excerpts default off), --clean, --base-title T, --allow-findings, --render-markdown, --max-items-per-page, --no-owners, --check, --with-eval[=FILE], --with-usage[=FILE].
A published catalog exposes names, descriptions, owners, token costs and lint findings: treat it like the configuration directory it describes.
Browser tests¶
go test ./internal/catalogsite also runs the site in a real headless Chrome or Chromium when one is installed
(AI_RULEZ_CHROME names the executable; -short skips them; they skip on their own when no browser is found or it
does not start). The harness speaks the DevTools protocol over --remote-debugging-pipe, so it needs no module and no
WebSocket. It opens every page of a rich and of a hostile fixture site from file:// and requires zero
Content-Security-Policy violations (recorded by a securitypolicyviolation listener the protocol injects), no console
error or warning, no uncaught exception, no JavaScript dialog, no inline script or style, and no request outside the
file system. It then filters, pages and navigates (overview, item, graph, MCP pages), and a control test injects an
inline script to prove the policy blocks it, so a zero count is not vacuous. The client itself is also tested against a
fake browser on every machine.
Design decisions¶
- Builder stays in
internal/govview. The issue namesinternal/catalog; the builder already lives ininternal/govviewnext to the roles and lock views that the CLI and the MCP tools share, and renaming would touch every caller for no behaviour change. The renderer is a separate package,internal/catalogsite. - Dual emit. Version 1 is the default and
--schema-version 2opts in (the issue's proposed answer). The HTML command always builds version 2. - Excerpt default. On locally, off with
--indexable(the issue's proposal);--include-excerptoverrides. - Secret check reuses
AR001. No new rule code. - Lint at the configured level. Findings come from the in-process lint engine with the project's
[lint]settings, not the strict level. - No
catalog-data.js. The pages are server-rendered and the filter reads the table rows, so a second copy of the data is not needed;catalog.jsonis the machine contract. - Approval. The overview has an Approval column (the status, or
not required); the item page shows the status with(required), the reviewers, their assurance level and the expiry, escaped like every other value. - Pagination is presentational. All rows are in
index.htmlin groups ofmax_items_per_page; the script pages them. Real per-page files would break filtering across pages and find-in-page.
Not yet built¶
--no-lint-messages, --link-sources, --single-file, a catalog diff page and approval or signature attestations
in the diff remain from the design.