MVP · measured 2026-09-21 · every number is from the project's records, listed at the end
seq 1Evalgate mvp · 2026-09-19 → 21
The MVP: fifty-three questions, held to account.
Forty-three packages in five days; 41 merged. Twenty campaigns of adjacent questions run live on 19 September — 53 of 53 finished, $8.27, on free search. Then every report was read, its defects packaged, and the four judged laws labelled. The tool got cheaper and more honest, and its judges were caught not judging. It did not yet get better at reading.
seq 2Runthe path · 17 → 21 Sep
The path.
The gate ran on 19 September from 09:46 to 11:05: twenty campaigns, one question after another, each under a cap the kernel refuses to cross. The next day three readers went through all 53 reports. The day after, the fixes merged — each with the reference tree re-recorded through free search.docs/GATE-MVP.md · docs/X2-REVIEW-2026-09-20.md · PRs #1164, #1123, #1190 · PR #1028 (BUILD-1: 226.7 s → 82.1 s)
seq 3Runrun.finished · 53
What it did.
Twenty campaigns, two to four questions each, the later ones meant to be answerable in part from the first's verified knowledge. Headless browsers, liveness checks, Postgres row-level security, idempotent retries, SQLite in production, Cloudflare Workers, drone rules, event sourcing, passkeys, email delivery. Every question ran once. Fifty-two finished; one ended as a research gap when the paid search month ran out at question 52.
That one gap is the most useful failure of the gate. A search outage was counted as a research outcome, which it was not. It became RETR-1 — five free search adapters, a per-engine quota book, a multi-engine seat, and the rule that a search outage ends the provider, not the run — and GOLD-SEAT, the reference tree recorded through free search. The MVP runs on no paid search.docs/GATE-MVP.md — Runs, Engines (S7) · docs/PACKAGES.json RETR-1, GOLD-SEAT · bench/sets/campaigns-mvp.json
reservations refused over cap · runs halted budget_exhausted
PASS
S6 · fetch ladder
163 / 203
failed fetches recovered · 40 refused · 0 silent
PASS
G2 · policy
1,984 / 1,984
fresh fetch and search rows carrying a policy version
PASS
E3 · six laws
106 / 106 · 120 / 212
L1, L3 by the ledger's rule · L2 L4 L5 L6 by two readers · judges not yet usable
MEASURED
B5 · knowledge
10 / 136
charter lines resolved from knowledge · 708 / 723 claims linked
MEASURED
rsk gate mvp report renders these from the ledger alone: a row the ledger cannot give says not measured and why, and a re-render of the same ledger is byte-identical. The gate measures that the tool ran as specified. It does not measure whether the answers are good — that is what the reading and the labels below are for.docs/GATE-MVP.md — Summary, Per campaign, Runs, Failed fetches (S6) · SPEC [C-284]–[C-289]
seq 5Decisionreview · labels · re-judge
What it taught.
The reading
Three independent readers, every report, against the question as asked, the pages the run opened, and the six laws. 264 defects, 65 of them high, in thirteen root causes. The kernel is cheap, honest about gaps, and quotes primary pages exactly when a search engine hands them over — Cloudflare's pricing to the cent, Transport Canada, sqlite.org. It under-reads: pages picked by search rank, the first 12,000 characters of each, three verified claims per run, and a synthesist writing past its own verdict ledger.
The synthesis gate's first live pass refused 3 of the 8 reference answers — names on the cited page but outside the quote, one term no page spelled. The rule was tightened to whole words of the cited page and the model repaired the invented term: 8 of 8, $1.08. QUAL-A re-recorded the reference at $0.99, GOLD-SEAT at $1.02; each replays at 0 misses, 0 deviations.docs/X2-REVIEW-2026-09-20.md · docs/PACKAGES.json QUAL-A, QUAL-B, QUAL-READ, QUAL-REPORT · crates/kernel/tests/fixtures/gold-8/baseline.json
The labels
L1 and L3 are decided by the ledger's rule. The other four laws need a reader. On 21 September every report was read twice more — by a sceptic who assumes the report is hiding something, and by the practitioner who asked the question — each labelling L2, L4, L5 and L6 with the sentence that decided it; a third reader settled every disagreement. Six of 53 reports pass all four.
The tool answers the question asked (L6) and says what it could not find (L5). What it gets wrong is stating an inference as something read on a page (L4) and softening what its own quote says (L2) — the two things QUAL-A and QUAL-B address. The labels were written by a model, not the owner: Fable 5.1, two readers and a tie-break. The 20 % the readers disagreed on is the honest error bar on the labels; the owner's page holds every label with its reason and can overrule any of them.docs/GATE-MVP.md — E3 calibration · doctrine/eval/v1/judges.toml [laws] mechanical, [calibration] · labelling workflow wf_d406ea4b-567
The judges
Two model judges asked the same four questions of every report, so that the labels could calibrate them and the cheaper one could stand in for a reader from then on. Neither can, yet. The flash seat passes every report on three of four laws; the pro seat is lenient too and ran out of its $3 cap 43 questions short of the fifty the doctrine needs. Where the readers found 92 failures, the judges found 12 and 30.
The ledger is doing its job here: a judge is never trusted by being listed, and these rows say the binary rubric with a 400-token answer is not a judge. The readers' brief — quote the sentence that decides it — is what the judge prompt lacks. That is the next evaluation-doctrine change.docs/GATE-MVP.md — E3, the judged laws · doctrine/eval/v1/JUDGES.md, judges.toml · the E-3 rows: judge_brier, judge_kappa, judge_peer_ca
The eight questions, again
The same eight questions the POC was judged on ran inside the campaigns, and the same readers judged each MVP report against the old engine's on the POC's rubric. Better 3, not-worse 1, worse 4. The bar is still six.
Every loss has one shape, and it is the one the reading found: the old engine read the canonical page — the vendor's pricing, the RFC, the maintainer's document — and this one read what the search engine ranked. That is what QUAL-A's canonical sources, QUAL-READ's authority tiers and the multi-engine seat are for. The verdicts are eight Decision.feedback events on the gate's runs.docs/GATE-MVP.md — E1 re-judge · docs/E1-2026-09-16.md · web-research/runs/<id>/report.md
seq 6Planstage mvp → prod
What comes next.
this weekQUAL-READ merges (authority tiers, query-focused windows, PDFs read). The judge prompt gets the readers' brief and is re-run against the 212 labels; the gate is called on the report as rendered. docs/GATE-MVP.md · E3, E7, E4
thenThe eight questions run once more on the finished main, judged the same way — the number this page will carry is that one. Twenty external reference tasks (DeepResearch Bench II, DeepResearch Bench, Cochrane) scored by rsk eval ext. GOLD-EXT
decisionsThe multi-engine seat as the default (#1111). A Stack Exchange fetch rung for the community hosts still refused (#1109). A search engine's failure surfaced instead of hidden (#1163).
productionMulti-node sync, the transparency log, tiering to R2, cloud clients, identity, SLOs — 18 packages. docs/PACKAGES.json stage prod
next: the roadmap — every package by stage, what it needs and when it is done — or try it.