Evaluate Your Search
1. List what's actually indexed
chops-search docs
This prints every indexed document with its URL and chunk count. The URLs are what your expectations must match, so start here: a mistyped expectation reads as a ranking failure rather than a typo, which is a bad afternoon. If a page you expected is missing, check the indexing rules before blaming the ranker.
2. Write the query set
Create fixtures/queries.toml beside your chops-search.toml. Four kinds of
case, each testing a different failure mode:
# Rare literal terms. The keyword engine must carry this.
[[query]]
q = "acks_late visibility timeout"
expect = ["/labs/celery-task-loss/"]
kind = "exact"
# Shares no useful words with the answer. Semantic must carry it.
[[query]]
q = "what happens to queued jobs when the process dies"
expect = ["/labs/celery-task-loss/"]
kind = "paraphrase"
# Typing toward a known page.
[[query]]
q = "about"
expect = ["/about/"]
kind = "navigational"
# Nothing answers this. The engine must return empty, not noise.
[[query]]
q = "sourdough starter hydration"
expect = []
kind = "negative"
expect lists URLs, any of which counts as a correct top-1. When two pages
legitimately answer the same query (a project page and the blog post about
it), list both rather than forcing an arbitrary winner. The
queries.toml reference has the full format.
3. Run it
chops-search build # eval scores the artifacts on disk, so build first
chops-search eval
You get a scoring: line stating exactly which configuration is being
measured (the values baked into index.bin, plus any flags), a PASS/FAIL
line per case, recall@1 and recall@3 by kind, real bytes-per-query numbers,
and for each failure the top 3 it returned instead. The eval drives the
actual engine over the actual byte path (plan, range-fetch, ingest, search),
so a bug in the loading logic shows up as a recall drop instead of hiding.
4. Diagnose a miss
For any failing case:
chops-search query "the failing query"
This prints the evidence behind the ranking: how the query tokenized on both sides, per-term keyword scores with document frequencies and per-field term frequencies, best-chunk cosine per document, and each engine's contribution to the fused order. When a result looks wrong the answer is almost always visible in the chunk count or in which field the term turned up in.
chops-search eval --explain prints the same evidence for every failing case
in one run, on the exact engine and flags the pass used.
Three patterns account for most misses:
- The expectation is wrong. The "wrong" winner genuinely answers the query. Fix the fixture, not the engine.
- The page is content-thin. Two chunks against a rival's thirty means the semantic side barely has surface to match. Fix the content.
- A stopword dragged in a rival. Look for a discriminating word that matches no documents while a common one scores. That's a ranking issue worth filing.
5. Calibrate, if the walk says so
Every scoring knob is baked into index.bin, and calibrate walks each
one against your query set without a rebuild:
chops-search calibrate
For every knob you get a table, one row per value, with the current value
marked > and every case that moved named under its row, then a verdict.
Nearly every verdict will be keep, stating the plateau the value sits on;
that is the answer, not a lack of one. A REVIEW names a value that gained
at least two cases, lists what it gained and lost, and re-runs it against
fixtures/known-failures.toml to name what it would break there. A named
casualty ends the review.
The fusion pair is coupled, so its joint grid stays in eval:
chops-search eval --sweep-rrf-k 2,4,8,16,32,60 --sweep-rrf-alpha 0,0.5,1,2
The loop for any nominated value is: explain each named flip
(calibrate --explain prints them on the candidate's own engine state, or
use chops-search query), decide whether the mechanism is real or a
coincidence, then pin the value in chops-search.toml and rebuild. Only
the committed, rebuilt value reaches the browser; the
configuration reference covers which keys bake
where. Save the transcript (-O or --clipboard): a dated calibrate run is
the thing you diff the next one against.
6. Measure the network side
Recall is one half of the calibration; what a query costs the visitor is
the other, and prefix_rows is the knob that trades one against the
other. Run it against the same query set:
chops-search plan
chops-search plan --prefix-rows 512,1024,2048,4096
Read the coverage: line first: it says how big the eager prefix would
need to be to cover 90% of the rows your queries need. Then read the table
as a trade: the eager column is what every visitor downloads once, the
bytes/q column is what a cold query fetches. Expect the shipped value to
sit on a plateau; if the next doubling of eager payload buys a rounding
error in bytes per query, that is a keep, and it is worth writing down
with the date the same way a calibrate transcript is.
7. Gate it in CI
chops-search eval --fail-under 0.85
Exit is non-zero below the threshold. Set it just under your measured baseline: high enough to catch a regression, low enough that one flipped case in a small set doesn't fail the build. Then wire it into your deploy workflow so a new post that quietly wrecks ranking becomes a red build instead of a discovery three weeks later.