MVP · records to 2026-09-23 · every number is from the project's records, listed at the end
seq 1Argumentclaim.asserted · anchors[0]
Research you can hold to account.
Every deep-research tool can search, read and write. None can be held to what it wrote. Here every claim, model call, dollar and judgment is an event you can open, replay and check. The report is prose; the reason is a ledger.
A finding in a real report, and the page it came from. The model cited a URL and a quote; the kernel found the quote in the bytes it had fetched and stored the offsets. The POC gate re-read every resolved anchor from the store: 451 of 451 matched, 0 mismatches. The MVP kept the rule and ran it over twenty campaigns: 53 questions, 708 of 723 live claims linked into the Knowledge ledger, $8.27 in all.docs/GATE-POC.md check 4 · docs/SPEC-kernel.md §2 anchor resolution · the report: Gold-8 Verdicts, Q6
Not released. The MVP runs and the console is built; the repository is private. Ask for access and we send one message when a question of yours can be run and held to account.
between them: how it works — the ledgers as rails, append-only, kill -9, replay at $0, budget — and the roadmap, 55 of the 56 MVP packages in, extraction and assembly next.
seq 3Evalbench.result
The headline numbers, each with its file.
EXTERNAL · 23 Sep
questions, rubrics and items written by other people; no run counted read the source its own rubric came from
check 8 fixed by FIX-590 the same day · in CI since
FIXED
E1 · side-by-side
5 of 8
not-worse 5 · worse 3 · the bar was 6
MISSED BY ONE
The POC values are bench.result events whose verdict is recomputed from value and target when the report renders. The MVP gate is rsk gate mvp report over the X2 ledger; a re-render of the same ledger is byte-identical. The external ten are one bench.result row per rubric item. Neither document is edited by hand.docs/GOLD-EXT-CLEAN.md · docs/GATE-MVP.md · docs/X2-REVIEW-2026-09-20.md · docs/BENCH-POC.md · docs/GATE-POC.md · docs/E1-2026-09-16.md