MVP · records to 2026-09-23 · every number is from the project's records, listed at the end
seq 1Note2026-09-23
A model that cannot write, put to work measuring.
We reached for TypeSafe's Jev as a cheap extractor. It cannot produce text at all — only a
probability, a choice among options you supply, or a score on a scale you define.
So it could never have been our extractor, at any price. The useful question was not what it
can write but what decisions we are making blind, at a scale nobody would pay a writing model to touch.
seq 2Unitthe shape
One request carries one page against every question.
A judgment costs about a thousandth of a cent, and the page is read once however many things
you need to know about it. That shape is why 500 questions over a whole corpus came to thirteen cents.
seq 3Doctrinecalibration
First, prove the instrument.
Before trusting a number from it we asked it about facts that do not exist: eight real
rubric items and six fabricated ones of identical shape, put to the same 26 pages.
Real facts 0.84 – 0.98, invented ones of identical shape 0.01 – 0.10, nothing in between —
so the 0.70 threshold everything below is read against is not a judgement call, and the instrument is
calibrated rather than merely cheap. A cheap instrument that agrees with you is worse than no instrument:
an unfalsified seat would have returned a high probability for a study that was never written, and this one
does not. The control is the only reason the number below is quoted at all, and EXTRACT-1's done-when
requires it to be re-run beside every future funnel rather than assumed to hold.docs/JEV-MEASUREMENT.md § the control, run before anything was believed — 8 real rubric items of drb2-062 and 6 fabricated, the same 26 stored pages, $0.0140 · issue #1404 done-when
seq 4Argumentthe finding
Where the score was actually going.
Every information-recall item of ten external benchmark tasks — 500 of them, written by
other people — put to every page those runs had already stored.
207 facts were on disk and never became a claim: 30.8 points of a 100-point benchmark
against the 16.42 the tool scores. More than the whole score, already paid for, unread.
then the sharper one
We asked whether any single passage carried a complete fact.
Present in pieces, complete nowhere. So the problem was never reading more — it is
assembly, joining facts across passages into records, and that killed the fix we were about to build.
seq 5Budgetwhat it cost
The price of knowing.
The loop matters more than the total. Diagnosing a change used to cost a full benchmark pass;
the funnel is thirteen cents and twenty minutes, so a hypothesis can be killed before it is built. That is
the whole return — not a better answer, a cheaper way to find out we were wrong.
next: the MVP — what the tool scores, and how that is measured — or the roadmap.