research-stack · research.devclusterai.com
MVP · records to 2026-09-23 · every number is from the project's records, listed at the end
seq 1Argumentsynthesis

Three questions this is built for. And one it is not.

Whether this fits you is a question about the shape of your work, not about how it connects. So: four cases, each with what you would ask, what comes back, what it costs and what it will not do — including the one where our own measurement says do not bother yet. Then the one thing to do about it. How it plugs into an agent is further down, where it belongs.

seq 2Argumentsynthesis

Where it fits, and where it stops.

four questions a reader brings · what the ledger answers · whether it fits today three the tool is built for · one the measurement says to wait on CASE 1 · FITS “I have to defend this report.” report.md is prose you can send · ledger_query answers where each sentence came from doctrine law 3 · docs/SPEC-kernel.md §7 — anchors, the byte range, the settled cost CASE 2 · FITS “Is this specific claim true?” a verdict, the search that went looking for its refutation, and the quote behind both doctrine law 4 and the VERIFY step · claim_status · confirmed / contested / refuted / unsupported CASE 3 · FITS “What does the evidence say — and where is it thin?” “What we could not establish” is a section of every report, not an omission doctrine law 5 · gaps reported, never smoothed · a gap run finishes, it does not error CASE 4 · NOT YET “A comprehensive survey, in the shape I was given.” measured 16.42 overall, information recall 12.89, over 672 expert-written items docs/GOLD-EXT-CLEAN.md · the fix is filed: EXTRACT-1 #1404, ASSEMBLE-1 #1405 source: doctrine/v1/DOCTRINE.md laws 3, 4, 5 · docs/SPEC-kernel.md §7 — the report's fixed sections case 4: docs/GOLD-EXT-CLEAN.md — the clean ten, no run having read the source its own rubric came from

The box is blue where the tool is built for the shape of the question and red where a measurement says it is not. Case 4 is red because it was measured and fell short, not because it is hard.doctrine/v1/DOCTRINE.md · docs/SPEC-kernel.md §7 · docs/GOLD-EXT-CLEAN.md · issues #1404, #1405

seq 3Argumentsynthesis · report.md

“I have to defend this report.”

The question is not what the answer is. You already have an answer; what you do not have is where it came from and what was not checked. This is the case the ledger exists for: the report is what you send, and the ledger is what you open when someone asks.

  • you askYour question, at a depth you choose, through the MCP tool or the CLI. Nothing is asked of you beyond the question itself.
  • what comes backreport.md, prose you can send, with a meta line carrying the depth, the plan and doctrine hashes, the counts and the cost — and no timestamps at all, so the same ledger renders the same file. Behind it, a row for everything: each claim’s anchor (the fetch it resolved against and the byte range [start, end)), the fetch.done of that page with its stored body, every model call with its settled micro-dollars, every rejection with its code and its path, and any verdict you wrote yourself as a Decision.feedback.One SQL statement opens any of it: ledger_query{sql}, read-only, one statement, 1,000 rows by default.
  • what it costsThe MVP gate ran 53 questions for $8.271496 — 51 at standard depth, one deep, one scout. The mean question was $0.16, the dearest $0.23, the cheapest $0.03, and the median run took 93.6 s. At deep depth under plan v2 a run averaged $0.40; none of those ten reached its $2.00 cap.
  • what it will not doIt will not tell you a page said something it did not — a claim needs a verbatim quote and an anchor or it is not a claim. It will also not pretend it read everything: in the gate, 14 of 217 first-attempt fetch failures were silent — every one of them unsupported_content_type, almost all PDFs, dropped with no fallback trail at all. Those pages are absent from the report and absent from the ledger, and that is the honest limit of “every claim traces to a row”.

The silent fourteen are listed one by one in the gate’s own render, with their URLs — which is the point: the failure mode is in the record rather than in a footnote.docs/SPEC-kernel.md §7 — Report, ledger_query · docs/GATE-MVP.md — the per-run table and the S6 fallback trails, the per-question figures derived from it · docs/DEPTH-1-NOTES.md — the ten deep runs on plan v2

seq 4Argumentverify · claim_status

“Is this specific claim true?”

Adjudication, not survey. One statement, and what you want is a verdict you can act on with the disconfirming evidence attached. The audit-first plan is built for exactly this shape: it interrogates the question before it searches, and then it goes looking for the refutation on purpose.

  • you askThe claim, put as a question. Or, cheaper, audit_question alone — an audit-only run that returns the question restated, its premises with the load-bearing ones marked, the sub-questions and the brief, before a single query is issued.
  • what comes backThe VERIFY step takes each load-bearing claim and searches specifically for evidence that it is false. A claim that survives a genuine attempt to refute it is confirmed; one that nothing addresses is unsupported, never confirmed by default. The report’s Contested and refuted section carries each mark, its verdict, and both sides — contested items list their undecided attackers, refuted items their accepted ones — read from claim_status by the fold.The word contested never appears in a model-written payload; it is computed, not asserted.
  • what it costsAn audit-only pass is the cheapest thing here, and on a ledger that already admitted the same auditor input it is a memo hit with no model call at all — the gate recorded 5 memo hits against 531 paid calls. A full standard run sits at the $0.16 mean above.
  • what it will not doIt will not hand you a yes or a no. The vocabulary is four verdicts, and unsupported — nobody addressed this — is a common and deliberate answer. If what you need is a confident binary, you will find this pedantic, and it is pedantic on purpose.

doctrine/v1/DOCTRINE.md — law 4 and the VERIFY step, verbatim · docs/SPEC-kernel.md §7 — audit_question, the Contested and refuted section, claim_status · docs/GATE-MVP.md — 531 paid calls, 5 memo hits

seq 5Argumentgaps · what we could not establish

“What does the evidence say — including where it is thin?”

The open question, where the useful answer may be nobody has established this. Most tools are built to avoid that sentence; here it has a section and a law behind it — gaps get reported, not smoothed.

  • you askA question you are genuinely unsure has an answer — including one whose premise may be wrong. The audit names the premises that are load-bearing before anything is searched.
  • what comes backFixed sections, in order, whether or not they are comfortable: What we could not establish, Assumptions the question was resting on, Surprises found during orientation, and — only when the audit moved — The question you should be asking. Every claim carries its epistemic label: observed (read it on a page), inferred, assumed; assumed never wears the costume of observed.A run that ends in a gap is finished, not failed: the Answer renders as the single line No answer: …, exit code 2. One of the gate’s 53 questions ended that way and counted as finished.
  • what it costsThe same as case 1 — a gap costs what an answer costs, which is the only way a tool can be indifferent between them.
  • what it will not doIt cannot yet tell you whether a gap is in the literature or in its own reading, and we have measured that this matters. Put to the pages ten runs had already downloaded, 207 of 500 expert-written recall items had their material sitting on disk and were never turned into a claim. On those tasks most of what was missing was missing from the report, not from the evidence.

That measurement is the reason case 4 below is honest rather than modest — we know where the loss is.doctrine/v1/DOCTRINE.md — laws 4 and 5, the epistemic labels · docs/SPEC-kernel.md §7 — the report's sections and the gap render · docs/GATE-MVP.md — 53 finished, 52 done, 1 gap · docs/JEV-MEASUREMENT.md — 207 of 500 (the measurement)

seq 6Evalbench.result · gold-ext/v3

“A comprehensive survey, in the shape I was given.” — not yet.

A required shape: named tables, fixed sections, a checklist of entities that must all be covered. If that is your case today, do not ask for access. Here is why, in the only terms that should persuade you — a measurement someone else wrote the rubric for.

Ten DeepResearch Bench II tasks, 672 expert-written rubric items, scored under gold-ext/v3 with no run having read the source its own rubric was derived from — the fetch lane refuses those by work identity, and three fetches were refused live during the pass.

dimensionmeasured, the clean tenbest publishedwhat it means for this case
overall16.4245.40the whole rubric, 672 items over ten tasks
information recall12.8939.98how much of what was asked for is in the report at all — a survey is mostly this
analysis27.1649.85does it explain and weigh, or only list
presentation27.0489.16the named tables and fixed sections a required shape asks for — the widest gap of the four

The right-hand column is a different judge. The published table (arXiv:2601.08536v3) was scored by Gemini-2.5-Pro, and its repository later moved to GPT-5.5; ours is Gemini-3.7-flash under our own prompts. The dimensions are comparable in shape, not in calibration, so read the column as a direction and not as a scoreboard. Our own before-and-after comparisons are sound because both sides are scored the same way.docs/GOLD-EXT-CLEAN.md — the ten tasks, the per-task table, the leaderboard row and why it is a different judge

What would change our mind

Not more reading. Asked whether any single passage carried a complete rubric fact, real items scored 0.48–0.64 against fabricated controls at 0.04: a study’s name sits in one place, its country, design and sample size in another. The constraint is assembly, not reading. EXTRACT-1 (#1404) takes statements off the page; ASSEMBLE-1 (#1405) joins them into records.

When a re-run of the clean ten moves information recall off 12.89, this page changes and says so. Until then it says this.docs/JEV-MEASUREMENT.md — the complete-or-partial split, 680 windows, $0.0643 · issues #1404, #1405 · the roadmap

seq 7Decisionaccess.requested

If one of the first three is your case.

There is one thing to do, and it is small. Ask for access — an address, and, if you like, the question you would run first.

  • what access isOne message, when a question of yours can be run and held to account: the report, the ledger it was rendered from, and the one SQL statement that opens any sentence in it. If you left a question, that is the one we run first.
  • what it is notThere is no download, the repository is not public, and there is no hosted product to sign into. Nothing is sent but that one message — no newsletter, no sequence, no product announcement.
  • the question you leaveIt is read. “What a report must prove before I would act on it” is exactly the question the evaluation ledger asks, and a question of yours is a candidate for the next campaign set.
  • removalReply to that one message, or write to [email protected]. The record is deleted, not flagged.

If case 4 is your case, the useful thing is not a request — it is the roadmap, and coming back when EXTRACT-1 and ASSEMBLE-1 have been measured.site/web/functions/api/waitlist.js — one KV record per sign-up · docs/PACKAGES.json — what is left before release

seq 8Runrun.finished · state

Where the project actually stands.

status · 2026-09-23

MVP. Fifty-five of the 56 MVP packages are in — ONBOARD-1 is the one still open, and the six-package BUILD lane is counted apart. The gate ran 53 of 53 questions over 20 campaigns for $8.27, on free search. Every report was read and its causes packaged; all four fix packages are merged. Against ten external tasks whose questions and rubrics other people wrote it scores 16.42 over 672 items, with no run having read the source its own rubric came from. Six of 53 reports pass all four judged laws. On the gate’s re-judge of the eight POC questions, 3 are better and 1 not-worse against the old engine — 4 of 8, where the first judging had 5 of 8. Since then: a console over the ledger, an article the writer must cite per paragraph, a blocklist enforced at the fetch lane, and the Jev measurement that redirected the next month of work. Not yet released: no download, the repository is not public.

The package count is derived, not remembered: stage mvp in docs/PACKAGES.json less the BUILD lane, against git log --merges main for pkg/<id>, with the umbrellas and the packages that landed inside another’s PR named as exceptions on the roadmap.docs/GATE-MVP.md · docs/GOLD-EXT-CLEAN.md · docs/JEV-MEASUREMENT.md · docs/X2-REVIEW-2026-09-20.md · docs/PACKAGES.json · docs/E1-2026-09-16.md · README.md

seq 9Unitunit.invoked · MCP

How it connects — the part that matters after the above.

Claude Code · OpenCode the researcher's agent MCP rsk mcp one config line audit_question deep_research_start / _status / _result deep_research · list_research_jobs research_doctrine · ledger_query ledger.sqlite3 + the report report.md → anyone asked “where did that come from?” ledger_query → which page, which bytes, what it cost, who decided what first: the technical researcher who can read a ledger · next: everyone who has to defend a report

An agent in Claude Code or OpenCode talks to the kernel over MCP. The ledger yields two things: report.md for anyone who is asked “where did that come from?”, and ledger_query for the researcher — which page, which bytes, what it cost, who decided what.docs/SPEC-kernel.md §7 interfaces

seq 10Unitunit.invoked · MCP

Inside Claude Code or OpenCode: rsk mcp.

rsk mcp speaks MCP over stdio with the same tool names as the server it replaces, so an existing mcp.web-research entry needs only its command changed. Legacy parameters are declared and ignored; tool failures come back as results with isError, never protocol errors.

The eight tools, and one prompt

  • audit_question(question)An audit-only run. Answers {job_id, run_id, question, restated, premise[]{text, load_bearing}, sub_questions[], brief} — the Question events of that run; on a ledger where the same auditor input was already admitted it is a memo hit, no model call.
  • deep_research_start(question, depth)depth = scout | standard | deep | exhaustive → {job_id, status, depth, note}. Spawns a detached worker (rsk research run <job_id>) so the client’s SIGTERM does not kill the run; the server only reads the ledger.
  • deep_research_status(job_id){job_id, question, depth, status: running | done | error, phase, message, elapsed_seconds, events[≤ 40]} — kernel states map onto phase; run.finished{state} onto status; a gap is done with a gap report.
  • deep_research_result(job_id, format)format = markdown | json.
  • deep_research(question, depth, wait_seconds)The blocking form, depth = standard, wait_seconds = 600 (clamped to 1,800). Progress notifications on every state.entered and at least every 30 s; on cancellation it returns {job_id, status: running}.
  • list_research_jobs(limit)limit = 20. Not list_jobs.
  • research_doctrine()No arguments; the doctrine.
  • ledger_query{sql, params?, max_rows?}→ {columns[], rows[][], truncated}. One read-only statement on the reader pool; DML, DDL, PRAGMA, ATTACH and BEGIN are refused; 1,000 rows or 1 MiB by default, hard cap 10,000; BLOBs base64 up to 4 KiB.
  • prompt · research(topic)The one MCP prompt the server declares.

docs/SPEC-kernel.md §7 — MCP (stdio, rmcp 3.3.0)

seq 11Runstate.entered · console

Not an MCP line: the console.

The same ledger has a front door for a person. rsk console serve puts five screens over it — ask (question, depth, and the cap it will run under, shown before anything runs), the run drawn live as its events land with the money counter beside it, the report byte-identical to rsk research result with every quote opening its page at the recorded bytes, the runs with cost, time and state, and the decisions — the kill switch, lanes and quotas, each written as a Decision through the owner path.

docs/CONSOLE.md · SPEC erratum C-350 … C-358 · crates/rs-console — beside your ledger today; the cloud seat beside rs-ingest reads the shared Postgres through the same reader (the rows materialised into a scratch ledger by the kernel's own triggers and fold) and waits on the owner's deploy

seq 12Argumentthe report, rendered

The report, section by section.

report.md is a pure function of the ledger: sections ordered by stable keys, never by time, and no timestamps — the meta line carries depth, the plan and doctrine hashes, counts and cost; wall time goes to report.json. A gap renders its Answer as No answer: …, a halted run as Halted: ….

# {question}
meta · depth · plan hash · doctrine hash · counts · cost
## The question you should be asking      (only when it differs)
## Answer                                  _Confidence: …_
## Findings
## Contested and refuted                   ### {mark} {claim} · **Verdict: …** · supporting · disconfirming
## Adversarial verification                confirmed · unsupported
## What we could not establish
## Next steps
## Assumptions the question was resting on
## Surprises found during orientation
## Sources                                 [host-minus-www](url)

The exact headers, in order — cases 1, 2 and 3 above are these sections read from the reader's side. Contested items list their undecided attackers and refuted items their accepted attackers, both read from claim_status; the word contested never appears in a model-written payload.docs/SPEC-kernel.md §7 — Report

And, under plan v2, an article

NARR-1 added a writer unit and a write state. In go the question, the coverage contract and the accepted claims with their anchors — never the contested or refuted — and out comes an article by facet with its own “what is not established”. Every paragraph carries at least one claim reference, and every number, date and name in it must be entailed by a cited quote or its page; an unanchored assertion is invariant.A7 at the paragraph that carries it.

The gate is strict enough to hurt: in the first ten-run pass three runs ended gap{invariant.A7} because the writer could not cite precisely enough on four words, discarding a finished synthesis. That is filed as #1326 and folded into EXTRACT-1 — a failed article should leave the run done with its report and the refusal named, not throw the research away.docs/PACKAGES.json NARR-1 · docs/DEPTH-1-NOTES.md — the merged plan v2 · issues #1284, #1326 · PR #1306

seq 13Fetchfetch.done · ledger_query

The ledger, from the command line.

Everything the report rests on is one SQL statement away: a claim’s anchors and byte range, the fetch.done row and stored body of that page, every model call with its settled micro-dollars, every rejection with its code and path. Nothing has to be trusted that can be opened.

  • rsk research start [--depth] [--run <id>]An existing run id resumes from the last admitted event.
  • rsk research status | result [--json]
  • rsk research replayRe-executes the plan at concurrency 1 with the model client and the retriever replaced by ledger-backed ones that refuse the network; any miss halts naming both keys.
  • rsk rebuild [--verify] [--out]Replays the events into a fresh file with the triggers firing, then compares view hashes; --verify prints table, rows before, after, equal and never swaps.
  • rsk bench <id>… · bench reportWrites bench.result events into the ledger; bench report recomputes every verdict and exits 1 on any gated FAIL.
  • rsk doctorDB path, node id, HLC skew, WAL size, sqlite version, the OpenRouter key and balance, the Brave key, and the six judged invariants printed pending.
  • rsk decide --scope <s> --text <t>Writes Decision.decided, actor owner.
  • rsk feedback --run … --target <id> --verdict <v> --reason <r>Writes Decision.feedback with refs [run, target], actor owner — the eight side-by-side verdicts were written this way.

Exit codes: 0 done, 2 gap, 3 halted, 1 error. Both Decision commands take the ledger lease before the kernel opens; a held lease is refused at once with nothing written.docs/SPEC-kernel.md §7 — CLI rsk

seq 14Runthe build

The build, as the README has it.

cargo build && cargo test          # offline; nothing in the suite opens a socket
cargo run --bin rsk -- doctor       # --db <path> for the ledger; exit 1 lists what is missing:
                                    # the OpenRouter key file or OPENROUTER_API_KEY, and BRAVE_API_KEY

A Rust workspace: rs (CLI), rsd (daemon: lanes, budget coordinator, fetch), rs-core, rs-store-sqlite, rs-fetch, rs-know and kernel — the kernel with its statechart runner, unit runner, MCP and the rsk CLI. Doctrine, plans, units and benchmark sets are versioned beside it. None of it is downloadable yet.README.md · docs/ADR-crates.md

if one of the first three cases is yours: ask for access · otherwise the roadmap, or what the MVP measured.