POC · measured 2026-09-16 · every number is from the project's records, listed at the end
seq 1Evalbench.result
Measured, not asserted.
Eleven micro-benchmarks, eight gate checks and an eight-question side-by-side, every value a bench.result event whose verdict is recomputed at render time; the gate document is generated, never edited. Seven of eight checks pass. Check eight fails, and stays failed here until the gate is re-run over FIX-590.
seq 2Evalbench.result · the line
The bench line: 47 gated rows, 0 FAIL.
K-1
0.34–1.70 ms
admission p99 per kind · target ≤ 2 ms
PASS
K-2
5,375–5,648 /s
sustained events · target ≥ 500
PASS
K-3
29 / 29 · 4.17 s
rebuild hash-equal @ 100k events · < 60 s
PASS
K-4
0.09 / 0.32 / 1.78 ms
anchor exact / nfc / folded @ 315 KB
PASS
K-5
37.3 MB
peak memory, deep run @ concurrency 8 · ≤ 60
PASS
K-6
0.002–0.009 ms
guard SQL p50 @ 10k–100k events · ≤ 10
PASS
A-6
11.8 ms
fold @ 1,000 claims, 50 % density · ≤ 50
PASS
U-4
78 / 78
replay served from the ledger · 0 calls
PASS
O-1
0 repeats
20 random kill -9 · report hash equal
PASS
G-2
0 over-cap
demand 400 vs cap 100 · 100 admitted, 30 refused
PASS
U-2
11 / 11
first-try valid, live, n = 20 · gated at n ≥ 200
RECORDED
targets from TESTPLAN §1.4 · release profile · mock model server · every verdict recomputed from value and target when the report renders
Measured 2026-09-16 on a MacBook Pro (macOS aarch64, 10 cpus, 32 GB), release profile, commit 79c3370, behind a load waiter so no sibling build could land mid-line. No gated row misses its POC target. U-2's live row is recorded from the gate ledger and not gated until n ≥ 200.docs/BENCH-POC.md · U-2 row: docs/GATE-POC.md (hits 11, total 11, memo_hits 9)
seq 3Evalgate · SPEC §8
Eight checks. Seven pass. The eighth is drawn.
Each check is a bench.result row in the gate ledger; the gate document is generated from them. Check eight fails: one live run at concurrency 6 cited a page a sibling task had fetched, and the same output is rejected when replayed at concurrency 1. FIX-590 is packaged against exactly that; the record on this page is the one measured before it.docs/GATE-POC.md (measured to 2026-09-16T16:12:59Z, commits 24bee4e and e41f64a) · docs/PACKAGES.json FIX-590
#
check (SPEC §8)
recorded
target
verdict
1
8/8 gold-8 headless at the baseline’s terminal state
docs/GATE-POC.md “The checks” — the recorded numbers column, condensed; the row ids and evidence fields are in the document
seq 4Decisionfeedback · actor owner
Eight questions, one judge: 5 of 8.
The old tool's report beside the kernel's for each of the eight benchmark questions, judged with the method written down first. Not-worse on five, worse on three; the bar was six. The kernel is far cheaper partly because it drives no browser — and that is where it loses.docs/E1-2026-09-16.md · bench/sets/gold-8/questions.json (old calls, cost, state) · kernel meta lines: Gold-8 Verdicts
The eight verdicts, after the self-check
Q1 · not-worse55/45 toward worse. Kernel structurally stronger — adversarial verification, honest medium confidence, no answer-vs-refutation contradiction — but never states the $0.02/browser-hour rate although it read the pricing page, gives no break-even threshold, and dates Browserbase agent runs to late 2026. Old answers decisively but contradicts its own refutation and lists one source as both supporting and disconfirming. Different failure modes, roughly even.headless browser TCO · deep
Q2 · worseThe question asks for the specific passive and active checks. Old maps them from primary vendor docs (FaceTec, Yoti, Jumio, Persona). Kernel answers with probabilistic ML ensembles, never addresses head pose, 3D SfM, frequency-domain or rPPG, reuses a fingerprint quote for a face claim, and its sources are mostly marketing blogs — 38 % primary against 76 %.selfie / liveness verification
Q3 · not-worseArguably better. Kernel gives a correct, direct two-sentence answer and two correct surprises (owner bypass; FORCE RLS does not touch superusers or BYPASSRLS). Old’s Answer says research did not run, then lists four richer surprises and four primary sources — self-contradictory.RLS as a backstop · scout
Q4 · not-worseBoth strong, different coverage. Old more primary and rigorous (rls.c, four CVEs, the covert channel, LEAKPROOF, PgBouncer). Kernel adds two real pitfalls old missed but repeats the superuser bypass three times, is blog-heavy (20 % primary) and says pg_dump may silently omit data when its own quote says row_security=off errors.RLS superuser / owner, incidents and practice
Q5 · worseOld anchored in canonical sources (Brooker, Google SRE ch. 22, the IETF draft, Stripe, Envoy) with three confirmed adversarial checks. Kernel’s tutorial-design reading is interesting but 11 % primary; the check-then-act race central to its Answer is marked unsupported by its own verification; confidence high anyway.idempotency and safe retries tutorial
Q6 · worseOld is a complete reference: jitter, per-try vs route timeouts, gRPC throttling, retry budgets, retriability, Retry-After. Kernel covers jitter and the 409/422 state machine from primary sources but omits retry budgets, timeouts, gRPC and retriability, says the server must return 422/409 while its own quotes say SHOULD, and pads sources with four versions of the same draft.idempotency retries backoff guide
Q7 · not-worseOld produced nothing (OpenRouter 402). Kernel delivers a coherent thesis — SLO multi-burn-rate alerting, high-cardinality events, static thresholds as the noise source — with one confirmed adversarial check, but has two authoritative sources among ~20 vendor blogs and never reaches pipeline-specific observability. Not-worse than nothing; not a clear pass against a good answer.observability and alerting
Q8 · not-worseSame claims and sources as Q7 served from cache (2 calls, $0.04), reordered findings, a slightly different Answer paragraph; substance and gaps identical. Verdict matches Q7.same question, cached re-run
The first draft had Q7 and Q8 as better; the self-check moved them to not-worse because the question's own bar is authoritative sources. The verdicts are recorded as eight Decision.feedback events, actor owner, in the gate ledger; the document is the reasoning they cite.docs/E1-2026-09-16.md — Verdicts, Tally
seq 5Fetchfetch.done · status failed
Reasoning ahead, sourcing behind.
The judge's finding that matters more than the tally: the kernel's discipline is ahead, its source acquisition behind. Primary-source share roughly halves wherever the old tool found primary sources, because the interim fetcher is locked out of the sites practitioners write on. That layer is what the substrate replaces.docs/E1-2026-09-16.md method steps 2 and 6, the finding, the ten defects
The ten defects filed from the side-by-side
defect 1Pricing page fetched ok but the headline rate not extracted or used (Q1) — unit / extraction.
defect 2StackOverflow 403 challenge on every URL — no challenge handling (S3).
defect 3Reddit robots-disallowed — an owner robots-policy decision, or accept.
defect 4No primary-source preference in ranking / selection — doctrine / search unit.
defect 5Cross-modality quote reuse: anchors byte-match but do not support the claim (Q2) — E3 judge.
defect 6Duplicate findings within a report (Q4, Q7) — fold / report dedup.
defect 7Date hallucination (late 2026, Q1) — date awareness in prompts.
defect 8A source marked both contested and confirming (Q5) — verification consistency lint.
defect 9Claim stronger than its quote: must vs SHOULD (Q6) — claim-strength lint / judge.
defect 10Claim contradicting its own quote (pg_dump silently omits, Q4) — adversarial pass miss.
Filed as MVP findings; the method that found them becomes the first evaluation doctrine (E7), and the same eight runs are to be re-judged after S1, through E3.docs/E1-2026-09-16.md “Defects to file” · docs/PACKAGES.json E7
next: the roadmap — FIX-590, then the substrate, identity, evaluation, driven mode and budget packages of the MVP.