research-stack · research.devclusterai.com
MVP · records to 2026-09-23 · every number is from the project's records, listed at the end
seq 1Argumentclaim.asserted · anchors[0]

Research you can hold to account.

Every deep-research tool can search, read and write. None can be held to what it wrote. Here the report is prose and the reason is a ledger you can open, replay and check.

THE PROBLEM
0 of 5
questions a report can answer about itself: which page said this · run it again · what did it cost · this is wrong · did anything fail
NO RECORD KEPT
WHAT IT DOES
13 ledgers
one append-only table. Every search, fetch, claim, call and dollar is a row; a quote not in the stored bytes is refused at commit, not flagged after
ON MAIN
WHAT IT COSTS
$0.156
a question — 53 for $8.271496, p50 94 s. Replaying a finished run: 79 of 79 calls served from the ledger, $0
MEASURED
WHAT IS PROVEN
490 / 490
resolved anchors re-read byte-for-byte at the POC gate, 0 mismatches, and the rule held over 53 MVP runs. Other people's rubrics: 16.42 of 100, behind the field
MEASURED

Each is a row a program rendered, not a figure anyone typed. Where a number is not clean, the page that carries it says so.the five questions: /problem · docs/SPEC-kernel.md §1 (the ledgers), invariant A1 · docs/GATE-MVP.md ($8.271496 over 53, p50 93.6 s) · docs/GATE-POC.md checks 3 and 4 · docs/GOLD-EXT-CLEAN.md (16.42 over 672 items)

report.md · Findings Stripe caches status codes and response bodies for keyed requests once endpoint execution begins, replaying even 500 server errors on subsequent retries. high confidence · docs.stripe.com “Stripe’s idempotency works by saving the resulting status code and body of the first request made for any given idempotency key, regardless of whether it succeeds or fails. Subsequent requests with the same key return the same result, including 500 errors.” A1 fetch.done · status ok · docs.stripe.com bodies/<sha256_text>.zst · extracted text, bytes Stripe’s idempotency works by saving the resulting status code and body of the first request made for any given idempotency key, regardless of whether it succeeds or fails. Subsequent requests with the same key return the same result, including 500 errors. start end the model cites url + quote · the kernel resolves fetch_ref and [start, end) · rungs exact → nfc → folded GATE-4 · every resolved anchor decoded from the store and compared byte-for-byte 490 / 490 · 0 mismatches

A quote is not checked after the fact — it is found in the bytes already fetched, or the claim is refused at commit. The gate re-read every resolved anchor from the store, 490 of 490, 0 mismatches, and the MVP held the rule over twenty campaigns.docs/GATE-POC.md check 4 · docs/SPEC-kernel.md §2 anchor resolution · the report: Gold-8 Verdicts, Q6

seq 2Planfour doors
seq 3Evalbench.result

The headline numbers, each with its file.

EXTERNAL · 23 Sep
questions, rubrics and items written by other people; no run counted read the source its own rubric came from
GOLD-EXT · the clean ten
16.42
672 rubric items · recall 12.89 · analysis 27.16 · presentation 27.04
RECORDED
JEV · the funnel
207 / 500
recall items whose material was already on disk · 30.8 points, unread
RECORDED
JEV · spend
$0.2484
every judgment of that measurement, nine runs · output free
RECORDED
MVP · 19–21 Sep
rendered from the gate ledger by a program; a row it cannot give says not measured
GATE-MVP · runs
53 / 53 · $8.27
20 campaigns · 52 done · 1 gap · 531 paid calls · p50 93.6 s
MEASURED
E3 · six laws
106 / 106 · 120 / 212
L1, L3 by the ledger's rule · L2 L4 L5 L6 labelled: 6 of 53 pass all four
LABELLED
E1 · re-judge
4 of 8
better 3 · not-worse 1 · worse 4 · the bar is 6
MISSED BY TWO
POC · 16 Sep
the gate as its last pass of that day recorded it, after FIX-590
GATE · check 4
490 / 490
resolved anchors re-read byte-for-byte · 0 mismatches
PASS
GATE · 8 checks
8 of 8
check 8 failed at 16:12, FIX-590 merged, re-measured at 22:49 · in CI since
PASS
E1 · side-by-side
5 of 8
not-worse 5 · worse 3 · the bar was 6
MISSED BY ONE

Nothing here is edited by hand: every verdict is recomputed from value and target as the report renders, and a re-render of the same ledger is byte-identical.docs/GOLD-EXT-CLEAN.md · docs/JEV-MEASUREMENT.md · docs/GATE-MVP.md · docs/X2-REVIEW-2026-09-20.md · docs/BENCH-POC.md · docs/GATE-POC.md · docs/E1-2026-09-16.md