Configuration

Configuration lives in chops-search.toml at the site root. Discovery walks up from the current directory the way cargo finds Cargo.toml, so commands work from anywhere inside the site. Paths resolve relative to the config file, not the working directory.

Every key is optional; an empty file (or no file) means the compiled defaults. An unknown key is an error, not a shrug: a misspelled min_gp = 0.08 would otherwise silently ship the corroboration gate disarmed, which is exactly the honesty gap the scoring keys exist to close.

chops-search.toml
content = "content"            # the Zola content tree to index
out     = "static/search"      # where artifacts + runtime land
model   = ".chops-search/model"

# dims = 128        # PCA target; unset means the model's native size (256)
# chunk_chars = 600
# prefix_rows = 2048

# BM25F field weights (body is fixed at 1.0)
# title_weight = 2
# tag_weight   = 4
# desc_weight  = 1

# Scoring calibration: baked into index.bin at build time
# min_gap   = 0.08   # corroboration gate; 0 (the default) disarms it
# rrf_alpha = 1.0    # confidence-weighted fusion; 0 (the default) is plain RRF
# min_cos   = 0.34   # relevance-floor OVERRIDE; unset derives from dims
# chunk_penalty = 0.05  # chunk-count correction; 0 disables

Shape keys

KeyDefaultNotes
contentcontentWalked recursively for .md files
outstatic/searchInside static/ so your SSG copies it verbatim
model.chops-search/modelThe lockfile lives beside this directory
dimsunset (native size)Real PCA, not truncation; init scaffolds 128. Re-eval after changing
chunk_chars600Smaller sharpens rare-word signal, costs more vectors. Values below 100 are rejected
prefix_rows2048Larger trades eager payload for fewer range requests. chops-search plan --prefix-rows measures the trade against your query set without a rebuild

BM25F field weights

What a length-normalised occurrence in each field is worth against one in the body (body is fixed at 1.0, so three knobs, not four). Each field's term frequency is normalised by that field's own average length before the weight applies, and saturation happens once on the combined value, so a weight biases without inflating. The useful range is therefore much smaller than a plain multiplier's would be; 0 ignores a field entirely.

KeyDefaultNotes
title_weight2
tag_weight4Tags outweigh titles: they're the author's own summary
desc_weight1Front-matter description, indexed as its own field. Parity with body is deliberate; it measured no better weighted up

Scoring calibration

min_gap, rrf_alpha, and min_cos are per-corpus calibrated values, and they follow the same provenance rule as the field weights: a value calibrated against a corpus travels with the corpus. build writes all of them into index.bin, and the engine reads them at construction, so the browser, a bare chops-search eval, and CI all score the same configuration from the same bytes. The eval and query flags still override per run for sweeping: the config states what ships, the flags state deviations from it.

KeyDefaultNotes
min_gap0 (gate disarmed)The corroboration gate threshold. When a query has no keyword evidence and no document stands out from the corpus median by at least this much, the semantic ranking is suppressed rather than served
rrf_alpha0 (plain RRF)Confidence-weighted fusion: the keyword list's vote scales by 1 + alpha × keyword confidence
min_cosunset (derived from dims)Relevance-floor override. Unset means the engine derives the floor from dimensionality, which tracks dims changes automatically and is right for almost every corpus. An explicit 0.0 is itself an override, meaning "floor off"
chunk_penalty0.02 (compiled)Chunk-count correction: coeff × √(2 ln n) subtracted from a document's best-chunk cosine, so a page with more chunks doesn't win on lottery tickets. 0 disables

min_gap, rrf_alpha, min_cos, and chunk_penalty are per-corpus calibrated values. These ship commented out in the scaffolded config because they are calibrated, not chosen: a value that helps one corpus hurts another. The loop is walk with chops-search calibrate, which reports plateaus and casualty-checked candidates per knob; verify the mechanism behind any candidate with calibrate --explain or chops-search query; then pin the value here and rebuild. See the evaluation tutorial and how ranking works.

Validation

Out-of-range values fail the build loudly rather than being clamped or ignored: min_gap, min_cos and chunk_penalty must be cosine-space values in 0..=1, rrf_alpha and the field weights must be finite and in 0..=100, and the build flags reject exactly what the file keys reject, so a flag cannot bake a value the config parser would have refused.

Precedence

Flag > file key > compiled default. The build flags (--dims, --chunk-chars, --prefix-rows, --content, --model, --out, --min-gap, --rrf-alpha, --min-cos, --chunk-penalty) override the file for a single run; the scoring flags on eval and query sweep values against a built index without rebuilding anything. Only what build bakes reaches the browser, and the build log's scoring: line states what shipped.