Skill Search¶
ai-rulez search ranks the skills your project serves against a query, with the same ranker as the find_skill
tool of mcp --serve-skills, and measures that ranking against labeled queries. By default it is
lexical and deterministic: no query or skill text is sent anywhere and the same input always gives the same order.
Optionally it is hybrid: the lexical list is fused with cosine similarity over embeddings you bring (an
OpenAI-compatible endpoint, Gemini through liter-llm, a local server, or a command), so a paraphrase such as
"customer wants money back" finds the skill described as "issue a refund".
- Searching
- The ranker
- Hybrid ranking
- Building the index
- Data egress
- Evaluating the ranking
- Cases file
- Result and gates
- Query mining
- Design decisions
Searching¶
ai-rulez search "customer wants money back"
ai-rulez search --format json --limit 10 deploy staging
ai-rulez search --role platform rotate credentials
ai-rulez search --mode hybrid --explain "customer wants money back"
The catalog is what the server would serve, selected with the same flags: --profile, --targets, --domain,
--allow, --deny, --source (repeatable), --role, --include-static, --offline and --frozen. --role and
--profile are mutually exclusive. Like the server, search (and search status) fetches the remote
skill sources the catalog is built from (--source, configured sources and includes) unless you pass --offline or
--frozen; that fetch is separate from the embedding egress described under Data egress. A skill the server refuses (lock or scan state) is not in the catalog and is
not ranked. Several words are joined into one query. --mode lexical|hybrid|vector overrides [search] mode
(and AI_RULEZ_SEARCH_MODE) for one run. A query that is the word index, status or mine alone is the
subcommand; add another word to search for it.
Text output lists rank, score, name, domain and description; --explain shows each skill's rank in the lexical
and the vector candidate list instead, and how the query was embedded. --format json prints
search.v1.schema.json:
{"schema_version": 1, "query": "money back", "ranking": "hybrid", "degraded": null, "count": 1,
"results": [{"name": "refund-policy", "description": "...", "score": 0.0325, "domain": "root", "digest": "sha256:...",
"lexical_rank": 3, "vector_rank": 1, "vector_sim": 0.71}]}
ranking is what produced the order. When a hybrid or vector search could not embed the query it ranks lexically
and degraded says why: no_index, provider_unavailable, timeout, budget or network_disabled; the reason
is also printed once on stderr. stale_vector: true marks a skill that changed since it was indexed: it ranks
lexically only. An empty or stopword-only query, or one nothing matches, returns no results in lexical mode. Unlike
find_skill, there is no in-role-first reordering: with --role the catalog already holds only that role's skills.
The ranker¶
Lexical: BM25F over four fields of a skill with weights name 3, triggers 2.5, keywords 2 and description 1, after
lowercasing, dropping stopwords and a light suffix stemmer (migrations matches migration). Hits are ordered by
score descending, then name. The code is internal/skillsearch; find_skill calls the same function, and a golden
test pins the order and scores of the lexical ranking so it cannot change by accident. It stays the reference and
the fallback of every other mode.
Hybrid ranking¶
Off by default. Turn it on with a [search] table and an embedding provider, then build the index once:
# .ai-rulez/config.toml
[search]
mode = "hybrid" # lexical (default) | hybrid | vector
fields = ["name", "triggers", "keywords", "description"] # the text sent to the embedder
index_body = false # true adds the first body_chars of SKILL.md
body_chars = 1200
fusion = "auto" # auto | rrf | weighted
rrf_k = 60
weights = { lexical = 1.0, vector = 1.0 }
vector_min_sim = 0.0 # abstain below this cosine (0 = off); see Abstaining
candidates = 50 # per list, before fusion
query_timeout_ms = 800 # the query embedding; on timeout the ranking is lexical
batch_size = 64 # texts per embedding call of `search index` (Gemini is sent one by one automatically)
index_dir = "local/search" # under the config dir; any directory outside local/ is meant to be committed
dtype = "float32" # float32 | float16 (half the size)
log_queries = false # user scope only (user config file or AI_RULEZ_SEARCH_LOG_QUERIES=1); see Query mining
# user config (~/.config/ai-rulez/config.toml): network settings only take effect here
[llm]
allow_network = true
provider = "gemini" # or base_url = "http://localhost:11434/v1" for Ollama, vLLM, LM Studio
embedding_model = "gemini-embedding-001"
api_key_env = "GEMINI_API_KEY"
max_calls = 200
The provider is the one [llm] table (see LLM access), so the network gate (allow_network, off by
default and user scope only), the budget (max_cost_usd, max_tokens, max_calls, fail closed), the response
cache and the redaction of provider errors all apply to every embedding call. [search.embeddings] model overrides
embedding_model for search only. AI_RULEZ_SEARCH_MODE overrides mode; a [search] table in the user config
overrides the repository's, key by key.
Fusion. Each list is cut to candidates and the two are combined:
auto(default): while every skill in scope has a current vector, rank by cosine alone (therankingfield then saysvector; an exact skill id in the query still pins first), because that beat every fusion on this repository's evaluation (below). As soon as one skill is new or changed since indexing, fall back torrfso that skill still ranks lexically. Setfusion = "rrf"or"weighted"to always fuse.rrf: reciprocal rank fusion,score = sum(weight / (rrf_k + rank))over the lists a skill is in. It needs no tuning of score ranges, because BM25 scores are unbounded and cosines are not comparable to them.weighted: each list's scores are min-max normalised to [0, 1] over its candidates and mixed witha = weights.lexical / (weights.lexical + weights.vector). More sensitive to the score distribution.
Why auto. Evaluated with Gemini gemini-embedding-001 on 36 labeled queries (plus 6 negatives) over this repository's
25 served skills, cut-off k = 3:
| Ranking | top-1 | hit@3 | MRR | nDCG@3 |
|---|---|---|---|---|
| lexical | 0.750 | 0.806 | 0.783 | 0.743 |
| hybrid, rrf (1:1, k = 60) | 0.778 | 0.861 | 0.833 | 0.792 |
| hybrid, rrf (1:2, k = 10) | 0.750 | 0.889 | 0.834 | 0.802 |
| hybrid, weighted (1:2) | 0.833 | 0.917 | 0.896 | 0.847 |
| hybrid, weighted (1:4) | 0.861 | 0.944 | 0.907 | 0.872 |
vector (also auto on a fresh index) |
0.861 | 0.972 | 0.910 | 0.894 |
No weight, rrf_k or fusion setting tried beat cosine alone, and only the vector ranking's paired gain over lexical
excluded 0 (MRR +0.126, 95% interval 0.007 to 0.261; the default RRF's was -0.004 to 0.110). Thirty-six queries is
small and the lexical-versus-paraphrase mix is this catalog's, so run search --eval --mode lexical,vector,hybrid
on your own cases before relying on it.
Ties order by name. A query that contains a skill's exact id as whole words ("use the deploy-staging skill") pins
that skill first. An id that is one ordinary word (test, build, fix) pins only when the query calls it a
skill ("use the build skill"): without that rule, "write the failing test first" pinned test over tdd-workflow
in an evaluation. Only exact ids pin, so a keyword-stuffed description gains nothing from it. vector mode ranks by
cosine alone (no id pin); a skill with no usable vector (new, or changed since indexing) follows the vector hits in
lexical order, and when no skill in scope has one the search is degraded: no_index. Unless you set vector_min_sim, a vector or hybrid result lists the nearest skills even when none is a good match,
because RRF scores have no absolute meaning (see Abstaining).
find_skill calls the same ranker. The MCP server never builds the index and never embeds a skill: it loads the
files search index wrote, reloads them when they change, and embeds only the query (bounded by
query_timeout_ms, with an in-memory cache of 256 queries on top of the [llm] cache). Its result gains ranking
and, when it fell back, degraded; see the MCP server.
Abstaining¶
A query nothing resembles ("what is the weather in Berlin") still has nearest skills. With vector_min_sim set, a
skill whose cosine to the query is below it is not a match: a vector ranking (including the auto default on a fresh
index) with none above it returns nothing, search --format json says "abstained": true and find_skill returns
no skill; rrf and weighted drop only the weak vector candidates, so a skill that matches by word still ranks. Skills without a current vector do not hide an abstention in vector mode. The user config's vector_min_sim wins over the repository's, and an explicit 0 there switches the threshold off.
Cosines differ per model, so no default ships: search --eval calibrates one from your cases. When the file has
positive and negative cases and a vector or hybrid mode ran, it prints the threshold that answers the most positives
while abstaining on the most negatives (calibration in the JSON) and each case's best cosine (top_sim). On this
repository's cases, vector_min_sim = 0.596 answered 36 of 36 positives and abstained on 5 of 6 negatives; the
near-miss "deploy the application to a kubernetes cluster" (0.619) still matched a skill. abstain_rate is the share
of negative cases answered with nothing. The advice comes from a small sample: re-run the evaluation with the value
set before committing it.
Building the index¶
ai-rulez search index --dry-run # the host, the number of texts and bytes, an estimate: nothing is sent
ai-rulez search index # embed what changed, write the index
ai-rulez search index --rebuild # re-embed everything
ai-rulez search index --items refund-policy,deploy-staging # force these, besides what changed
ai-rulez search status # state of the index against the served skills
$ ai-rulez search index --dry-run
provider gemini@default (allow_network=true)
model gemini-embedding-001
items 25 cached 0 to embed 25 withheld 0 est. tokens 1627 est. cost <= $0.000244
fields name, triggers, keywords, description body: no
egress 25 texts, 4.7 KiB -> default
The embedded text of a skill is a fixed template over the configured fields (name:, description:, triggers:,
keywords: lines, then body: when index_body is set). A vector is reused when the provider, the model, the
fields and the SHA-256 of the exact text match, so editing a skill's body (with index_body = false) costs nothing, one edited description costs one
text (so does a rename, since the name is part of the text unless fields leaves it out), and a changed model or field set re-embeds everything:
vectors of different models are never mixed. A batch the provider rejects (a 4xx, an over-long input) is split in halves
until the refused skill stands alone, which is then skipped and named on stderr (rejected in the JSON summary);
the run goes on and exits 2. Five rejections in a row stop the build, since that is a broken provider rather than
bad skills. A budget, network, key or transient provider error stops the build too. Either way the vectors that
finished are written (atomically: vectors.bin, then manifest.json), the skills still missing are printed and the
run exits 2; those rank lexically. A second concurrent run is refused by index.lock.
search status reads files only and reports none, unreadable, incompatible (other provider, model, fields or
template), stale or fresh, with the skills that are current, changed since indexing (the embedded text
differs), not indexed, or indexed but no longer served.
The index is two files in index_dir:
manifest.json: schema and text-template versions, provider fingerprint, model, dimensions, dtype, fields, the digest ofvectors.bin, and per skill its id, domain, lock digest and text digest. It is deterministic (sorted, no timestamps), so a committed index diffs cleanly and rebuilding from unchanged inputs rewrites nothing.vectors.bin: little-endian, row-major, L2-normalisedfloat32(orfloat16: 25 skills of 3072 dimensions are 300 KiB, or 150 KiB in half precision). Loading checks the length against the manifest, the SHA-256, and rejects NaN, infinities and rows that are not unit length (a scaled-up row in a committed index would otherwise win every query); a file that does not validate is "no index", never a crash. Search is a brute-force dot product, clamped to [-1, 1], fine to tens of thousands of skills.
The default local/search is machine-local and gitignored. Set index_dir to a directory outside local/ (for
example search-index) to commit the index so CI can run a hybrid evaluation without a key; validate then
checks it (AR9D1, a warning) when a description changed, the model changed or the files are missing or damaged,
and search status gives the full comparison against the served catalog. A poisoned index can only change the
order of results: load_skill still enforces the lock, approval and scan state of what it returns.
Command provider¶
A program can produce the vectors instead of [llm]:
# user config only (or --allow-exec with a repository config)
[search.embeddings]
command = ["/usr/local/bin/embed"]
model = "my-local-model" # recorded in the index
pass_env = ["EMBED_TOKEN"]
It runs as an argv (no shell) in the project root with only PATH, HOME and the pass_env variables, a 30 s
timeout and a 64 MiB stdout cap. It reads {"input": ["text", ...]} on stdin and prints
{"vectors": [[...], ...]}, one vector per input in order. Because it runs a program, a repository config cannot
set it: it is honoured from the user config file, or from the repository config when you pass --allow-exec
(otherwise it is ignored with a warning). The MCP server only honours the user config.
Data egress¶
| Operation | Data sent | To |
|---|---|---|
search index |
One text per changed skill: the configured fields (default name, triggers, keywords, description, already visible to agents) and, only with index_body, the first body_chars of the body |
the embedding endpoint |
search, find_skill in hybrid or vector mode |
The query text (capped at 2 KiB; an agent's task description can hold user data) | the same endpoint |
| lexical mode | nothing | nowhere |
Before sending, each skill text goes through the secret scanner of AR001; a hit withholds that skill
(AR9D3, reported on stderr) instead of sending a masked string, and the skill ranks lexically only. Queries are
not scanned. allow_network stays off until you turn it on in the user config, loopback hosts included, and
--dry-run shows the host first. Choose a local or in-region endpoint for residency: the provider fingerprint is
recorded in the manifest, so a committed index shows where its vectors came from.
Evaluating the ranking¶
ai-rulez search --eval search-cases.yaml
ai-rulez search --eval search-cases.yaml --min top1=0.6,mrr=0.7
ai-rulez search --eval search-cases.yaml --mode lexical,hybrid # side by side, with paired intervals
ai-rulez search --eval search-cases.yaml --output result.json
ai-rulez search --eval search-cases.yaml --baseline result.json --max-flips 0
ai-rulez search --from-evals # cases derived from the skills' eval cases
Use it to see whether a change to a skill description, to the ranker or to the embedding model helped, and to gate
a pull request on it. --eval takes the queries from the file; it does not accept a query argument. --k
overrides the file's cut-off; a value above 100 is an error, not clamped. --mode takes a comma-separated list
(lexical, vector, hybrid, or all); the first is the primary mode the gates check, and the default is the
configured mode. A mode that needs vectors needs the index.
--from-evals derives cases from the eval-runner cases under skills/<name>/evals/*.eval.yaml (see
Evals), alone or added to --eval: a prompt with expect_trigger: true for skill S becomes a case
that expects S; a near-miss prompt, or any prompt with expect_trigger: false, becomes a case with avoid: [S]
(S must not rank first). Their ids are <skill>/<case> and they are tagged from-evals. A case whose skill is not
in the served catalog is skipped and counted on stderr.
Cases file¶
version: 1 # required
k: 5 # cut-off of recall@k, hit@k and nDCG@k (default 5, max 100)
cases:
- id: refund-paraphrase
query: "customer wants money back for a double charge"
expect: [refund-policy] # skill names; any of them found counts
tags: [paraphrase, billing]
- id: graded
query: "set up staging deploy pipeline"
expect: [{id: deploy-staging, grade: 3}, {id: ci-pipeline, grade: 1}] # grades enable nDCG
- id: scoped
query: "rotate credentials"
role: platform # rank only the skills this role serves
expect: [rotate-keys]
- id: near-miss
query: "deploy to production"
expect: []
avoid: [deploy-staging] # must not rank first
- id: negative
query: "what is the weather tomorrow"
expect: [] # nothing relevant; reported separately
Skills are keyed by their catalog name. Problems in the file are reported together as AR9D2 (exit 1): an unknown
field, a missing version, id or query, a duplicate case id, a duplicate or unknown skill in expect or
avoid, a grade outside 1-9, a skill both expected and avoided, an unknown role, or an expected skill the role
does not serve. A bare name has grade 1.
Result and gates¶
Per mode: top1 (a relevant skill is first), recall_at_k (relevant skills in the top k over relevant skills),
hit_at_k (any relevant skill in the top k), mrr (1 over the rank of the first relevant skill, 0 if none),
ndcg_at_k when a case is graded (gain is the grade, discount log2(rank + 1), normalised by the ideal order of
that case's grades) and avoid_top1 when cases carry an avoid list (the share where an avoided skill ranked
first; lower is better), overall and per tag. Negative cases are left out of the metrics and listed with the skill
that ranked first and its score, which is the ranker's own (BM25F, RRF or cosine) and only informational. A 95%
bootstrap interval (1000 resamples, fixed seed) is printed for top1, recall_at_k, mrr and ndcg_at_k; it is
never gated on. With several modes, paired_vs_lexical gives each mode's mean difference to lexical over the same
cases with its own interval: an interval that excludes 0 is a real change, which two overlapping per-mode intervals
cannot tell you. A vector mode that fell back to lexical for any case (no network, budget, timeout) is an error
(exit 1), never a pass, because the numbers would not measure the ranker asked for.
--format json prints search-eval.v1.schema.json,
which is also what --output writes and a later --baseline reads. The top-level cases, tags, intervals, misses and
negatives are those of the primary mode; by_mode holds every mode. Each case records rank (null when not
retrieved) and hit.
| Gate | Fails when |
|---|---|
--min top1=0.6,recall=0.8,hit=0.9,mrr=0.7,ndcg=0.7 |
a metric of the primary mode is below its floor |
--baseline prev.json --max-flips N |
more than N cases went from hit to a miss since the baseline (default 0) |
With a small set one case moves top1 by several points, so prefer the flip count to a raw floor. Cases present in
only one of the two runs are not compared. A baseline must come from the same primary
mode and the same --k; otherwise the run is refused (exit 1). --output is not written when the embeddings were
degraded, so a fallback run cannot become a baseline.
Exit codes: 0 every gate passed, 1 the command could not run (bad flags, invalid cases file, unknown skill,
unreadable baseline, a mode without an index, degraded embeddings), 2 a gate failed (or search index stopped
early). Failed gates are printed to stderr as AR9D4.
Query mining¶
Real queries are the best test cases, but the usage log holds identifiers only. Mining is a
separate, opt-in step: with [search] log_queries = true in the user config file (or AI_RULEZ_SEARCH_LOG_QUERIES=1; a
repository config cannot turn it on; search and the MCP server warn when it tries), find_skill appends each query's text (secret-looking
queries are skipped, text is capped at 512 bytes) and, when the session then calls load_skill, the skill it
loaded, to <config dir>/local/search-queries.jsonl (mode 0600, gitignored, never sent anywhere, capped at 8 MiB,
session ids hashed). search mine turns it into cases:
ai-rulez search mine --output mined-cases.yaml # a query followed in its session by loading S expects S
ai-rulez search mine --min-count 2 --purge # keep queries seen twice with the same skill; delete the log
A query's label is the first skill loaded in its session before the next query; queries followed by different
skills with no clear winner, or by nothing, are dropped. The label is weak, since the agent chose the skill, which
does not make it right, so the cases are tagged mined and meant to be reviewed before they gate anything.
Design decisions¶
Follows issue #222; where it left a choice open the proposed default is taken.
- Default stays lexical, offline and deterministic. A hybrid or vector ranking always falls back to lexical
with a
degradedreason; it never fails a search or afind_skillcall. - Embeddings go through
internal/llm(one egress configuration, budget and gate), including liter-llm for Gemini and 170 other providers. The command provider is the escape hatch for anything else. - The index is keyed by the SHA-256 of the exact embedded text, not by the skill digest, so an edit that does not
change the embedded text is free.
search statusandstale_vectortherefore call a skill stale when its embedded text changed (the design said "item digest"); a changed digest with the same text is reported as drift and still ranks with its vector. - Staleness in
validate(AR9D1) compares the project's own skills present in both the index and the content tree; which skills an index holds depends on the serve flags it was built with, so missing and orphaned skills are reported bysearch statusonly. - Loopback endpoints get no exemption from
allow_network, and network settings are honoured from the user config only (see the trust rule in LLM access). dtype = "float16"and[search] log_queriesare additions to the design's config; calibration is part ofsearch --evalrather than asearch status --calibrateflag, and there is no LLM-suggested paraphrase step.- The cases file is YAML (as proposed in the issue) and read strictly: unknown fields are errors, so a file never
means less than it says. Flip comparison uses
hit_at_k, the issue's "hit to miss". AR9D2andAR9D4are emitted bysearch --eval,AR9D3bysearch index,AR9D0andAR9D1byvalidate.search_skills(the MCP listing tool) stays lexical.