Ledger Kernel · research-stack
MVP · measured 2026-09-21 · every number is from the project's records, listed at the end
seq 1Evalgate mvp · 2026-09-19 → 21

The MVP: fifty-three questions, held to account.

Forty-three packages in five days; 41 merged. Twenty campaigns of adjacent questions run live on 19 September — 53 of 53 finished, $8.27, on free search. Then every report was read, its defects packaged, and the four judged laws labelled. The tool got cheaper and more honest, and its judges were caught not judging. It did not yet get better at reading.

seq 2Runthe path · 17 → 21 Sep

The path.

17 Sep MVP begins 43 packages 18 Sep the build 227 s → 82 s 19 Sep the gate 53 runs · $8.27 20 Sep the reading 13 causes 21 Sep fixes merged SEAT · A · B 21 Sep labelled 212 · 8 re-judged

The gate ran on 19 September from 09:46 to 11:05: twenty campaigns, one question after another, each under a cap the kernel refuses to cross. The next day three readers went through all 53 reports. The day after, the fixes merged — each with the reference tree re-recorded through free search.docs/GATE-MVP.md · docs/X2-REVIEW-2026-09-20.md · PRs #1164, #1123, #1190 · PR #1028 (BUILD-1: 226.7 s → 82.1 s)

seq 3Runrun.finished · 53

What it did.

Twenty campaigns, two to four questions each, the later ones meant to be answerable in part from the first's verified knowledge. Headless browsers, liveness checks, Postgres row-level security, idempotent retries, SQLite in production, Cloudflare Workers, drone rules, event sourcing, passkeys, email delivery. Every question ran once. Fifty-two finished; one ended as a research gap when the paid search month ran out at question 52.

the campaign set · 20 campaigns · 53 questions · one square per run 01 02 03 04 05 06 07 08 09 10 11 12 13 14 15 16 17 18 19 20 gap done · 52 research gap · 1 — the search month ran out at question 52 the eight POC questions, asked again inside their campaigns c01 headless browsers · c02 liveness · c03 Postgres RLS · c04 idempotent retries · c05 pipeline alerting · c06 SQLite c07 async Rust · c08 LLM billing · c10 Workers · c12 drone BVLOS · c13 event sourcing · c15 passkeys · c17 email · c20 LLM judges

That one gap is the most useful failure of the gate. A search outage was counted as a research outcome, which it was not. It became RETR-1 — five free search adapters, a per-engine quota book, a multi-engine seat, and the rule that a search outage ends the provider, not the run — and GOLD-SEAT, the reference tree recorded through free search. The MVP runs on no paid search.docs/GATE-MVP.md — Runs, Engines (S7) · docs/PACKAGES.json RETR-1, GOLD-SEAT · bench/sets/campaigns-mvp.json

seq 4Evalgate mvp · summary

The numbers.

questions
53 / 53
20 campaigns · 52 done · 1 gap · 0 halted · 0 re-runs
MEASURED
spend
$8.27
531 paid calls · 5 memo hits · settled on the Budget ledger
MEASURED
wall · p50
93.6 s
run.started → run.finished · n = 53 · deep, standard, scout
MEASURED
G1 · budget
0 · 0
reservations refused over cap · runs halted budget_exhausted
PASS
S6 · fetch ladder
163 / 203
failed fetches recovered · 40 refused · 0 silent
PASS
G2 · policy
1,984 / 1,984
fresh fetch and search rows carrying a policy version
PASS
E3 · six laws
106 / 106 · 120 / 212
L1, L3 by the ledger's rule · L2 L4 L5 L6 by two readers · judges not yet usable
MEASURED
B5 · knowledge
10 / 136
charter lines resolved from knowledge · 708 / 723 claims linked
MEASURED

rsk gate mvp report renders these from the ledger alone: a row the ledger cannot give says not measured and why, and a re-render of the same ledger is byte-identical. The gate measures that the tool ran as specified. It does not measure whether the answers are good — that is what the reading and the labels below are for.docs/GATE-MVP.md — Summary, Per campaign, Runs, Failed fetches (S6) · SPEC [C-284]–[C-289]

seq 5Decisionreview · labels · re-judge

What it taught.

The reading

Three independent readers, every report, against the question as asked, the pages the run opened, and the six laws. 264 defects, 65 of them high, in thirteen root causes. The kernel is cheap, honest about gaps, and quotes primary pages exactly when a search engine hands them over — Cloudflare's pricing to the cent, Transport Canada, sqlite.org. It under-reads: pages picked by search rank, the first 12,000 characters of each, three verified claims per run, and a synthesist writing past its own verdict ledger.

53 reports read · 264 defects · 13 root causes · 4 packages QUAL-A merged 21 Sep #1 no canonical sources #6 no coverage contract #7 no run date #8 no arithmetic (½) QUAL-B merged 21 Sep #4 synthesis not gated #5 uncited surprises #8 no arithmetic (½) QUAL-READ merging #3 pages by search rank #9 first 12k chars only #10 PDF, Markdown unread QUAL-REPORT merged 21 Sep #2 verdicts unreconciled #11 renderer hygiene #13 reuse by term overlap open: #12 community and docs hosts refused at rung 0 — waits on #1109 (a Stack Exchange fetch rung) and on #1163 (a search engine's failure hidden as an empty result) each package landed with the reference tree re-recorded and replayed at 0 misses, 0 deviations

The synthesis gate's first live pass refused 3 of the 8 reference answers — names on the cited page but outside the quote, one term no page spelled. The rule was tightened to whole words of the cited page and the model repaired the invented term: 8 of 8, $1.08. QUAL-A re-recorded the reference at $0.99, GOLD-SEAT at $1.02; each replays at 0 misses, 0 deviations.docs/X2-REVIEW-2026-09-20.md · docs/PACKAGES.json QUAL-A, QUAL-B, QUAL-READ, QUAL-REPORT · crates/kernel/tests/fixtures/gold-8/baseline.json

The labels

L1 and L3 are decided by the ledger's rule. The other four laws need a reader. On 21 September every report was read twice more — by a sceptic who assumes the report is hiding something, and by the practitioner who asked the question — each labelling L2, L4, L5 and L6 with the sentence that decided it; a third reader settled every disagreement. Six of 53 reports pass all four.

the six laws · 53 runs · pass (blue) · fail (red) 53 L1 audit first 53 / 53 mechanical · the ledger's rule L3 source opened 53 / 53 mechanical · the ledger's rule L2 evidence outranks priors 18 / 53 labelled · 13 of 53 settled by a third reader L4 epistemics labelled 13 / 53 labelled · 18 settled by a third reader L5 gaps reported, not smoothed 38 / 53 labelled · 9 settled by a third reader L6 reframe offered, not imposed 51 / 53 labelled · 3 settled by a third reader L2 L4 L5 L6: two independent readers per report — a sceptic, and the practitioner who asked a third reader settled the 43 of 212 they disagreed on · reader agreement 80 %

The tool answers the question asked (L6) and says what it could not find (L5). What it gets wrong is stating an inference as something read on a page (L4) and softening what its own quote says (L2) — the two things QUAL-A and QUAL-B address. The labels were written by a model, not the owner: Fable 5.1, two readers and a tie-break. The 20 % the readers disagreed on is the honest error bar on the labels; the owner's page holds every label with its reason and can overrule any of them.docs/GATE-MVP.md — E3 calibration · doctrine/eval/v1/judges.toml [laws] mechanical, [calibration] · labelling workflow wf_d406ea4b-567

The judges

Two model judges asked the same four questions of every report, so that the labels could calibrate them and the cheaper one could stand in for a reader from then on. Neither can, yet. The flash seat passes every report on three of four laws; the pro seat is lenient too and ran out of its $3 cap 43 questions short of the fifty the doctrine needs. Where the readers found 92 failures, the judges found 12 and 30.

pass rate per law · readers (blue) · flash judge (grey) · pro judge (hatched) 100 % L2evidence outranks priors 34 % 100 % 87 % L4epistemics labelled 25 % 100 % 85 % L5gaps reported 72 % 100 % 88 % L6reframe offered 96 % 77 % 59 % flash: 199 pass of 212 · κ 0.00 on L2, L4 and L5 — it passes every run · Brier 0.60 / 0.69 / 0.26 / 0.16 pro: 153 of 212 answered before its $3 cap · κ 0.04–0.19 · under 50 queries, so not calibrated peer agreement between the two judges ≈ 0 · the doctrine calibrates a judge at 50 labelled queries (judges.toml)

The ledger is doing its job here: a judge is never trusted by being listed, and these rows say the binary rubric with a 400-token answer is not a judge. The readers' brief — quote the sentence that decides it — is what the judge prompt lacks. That is the next evaluation-doctrine change.docs/GATE-MVP.md — E3, the judged laws · doctrine/eval/v1/JUDGES.md, judges.toml · the E-3 rows: judge_brier, judge_kappa, judge_peer_ca

The eight questions, again

The same eight questions the POC was judged on ran inside the campaigns, and the same readers judged each MVP report against the old engine's on the POC's rubric. Better 3, not-worse 1, worse 4. The bar is still six.

the eight questions against the old engine · POC (16 Sep) · MVP (21 Sep) Q1 headless TCO not-worse worse Q2 liveness worse worse Q3 RLS scout not-worse better Q4 RLS owner not-worse not-worse Q5 idempotency worse worse Q6 retries worse worse Q7 alerting not-worse better Q8 alerting, re-run not-worse better POC5 of 8 MVP4 of 8 the same rubric, e1-verdict/v1 · the bar is 6 · Q7 and Q8: the old run failed on a 402 and produced no report Q1: the MVP never reached the pricing page the old engine read · Q3: the old report never answered

Every loss has one shape, and it is the one the reading found: the old engine read the canonical page — the vendor's pricing, the RFC, the maintainer's document — and this one read what the search engine ranked. That is what QUAL-A's canonical sources, QUAL-READ's authority tiers and the multi-engine seat are for. The verdicts are eight Decision.feedback events on the gate's runs.docs/GATE-MVP.md — E1 re-judge · docs/E1-2026-09-16.md · web-research/runs/<id>/report.md

seq 6Planstage mvp → prod

What comes next.

  • this weekQUAL-READ merges (authority tiers, query-focused windows, PDFs read). The judge prompt gets the readers' brief and is re-run against the 212 labels; the gate is called on the report as rendered. docs/GATE-MVP.md · E3, E7, E4
  • thenThe eight questions run once more on the finished main, judged the same way — the number this page will carry is that one. Twenty external reference tasks (DeepResearch Bench II, DeepResearch Bench, Cochrane) scored by rsk eval ext. GOLD-EXT
  • decisionsThe multi-engine seat as the default (#1111). A Stack Exchange fetch rung for the community hosts still refused (#1109). A search engine's failure surfaced instead of hidden (#1163).
  • productionMulti-node sync, the transparency log, tiering to R2, cloud clients, identity, SLOs — 18 packages. docs/PACKAGES.json stage prod

next: the roadmap — every package by stage, what it needs and when it is done — or try it.