research-stack · research.devclusterai.com
MVP · records to 2026-09-23 · every number is from the project's records, listed at the end
openfindingseach checked against the board · 2026-09-23

What is broken.

Six defects in what is already merged and published here. None of them is fixed. A project whose whole claim is that it records what it could not establish has to record this too.

seq 1Noteopen · 2026-09-23

Six defects a reader of this site can see.

what four open defects have already cost · checked 2026-09-23 the bar is the share of what each one touched · each row carries its own denominator #1326 finished runs discarded 3 / 10 runs the research finished and the synthesis is on the record; the article gate refused the writing #1327 a report that is a failure log 130 / 260 lines one plan-v2 report, counted by a blind reader who was not asked to look for it #1328 analysis lost, same 8 tasks 31.21 → 16.42 recall rose 99 % over the same eight tasks, so the overall score barely moved #1401 runs of ci on main, red 18 / 18 runs every run since 21 Sep 05:52 · one leg, macos-15, fails in about 7 s with no steps recorded #1330 is not here: the defect is fixed in the code and the issue is still open. #1398 has no number to draw. sources: issues #1326 #1327 #1328 #1401 · docs/DEPTH-1-NOTES.md · gh run list --workflow ci
  • #1326A run whose writer cannot get its article past the citation gate ends gap, and the finished research is discarded — 3 of 10 runs in the plan-v2 external pass, each with its synthesis already on the record. Being fixed inside EXTRACT-1: a refused article should leave the run done with its report and the refusal named in its meta line.open · NARR-1-F8 · docs/DEPTH-1-NOTES.md, the live pass
  • #1327The report pads. About 130 of one report’s 260 lines were a list of failed fetches, Findings repeated Sources verbatim, and “what is not established” was printed twice. Presentation scores 27.04 on the clean ten, and this is part of why.open · NARR-1-F9 · said unprompted by the blind pairwise readers
  • #1328Analysis fell 47 % — 31.21 to 16.42 over the same eight tasks — when the cover loop doubled what each run read. Our first explanation for it measured false: we said the synthesist was covering more claims under a fixed output budget, and the ledger says that budget never bound — 39 of 39 calls ended stop, the synthesist using at most 3,670 of its 8,000 tokens. The article is claim-starved instead, and 19 of 21 of the mechanism terms the judge missed were sitting in pages the run had already stored.open · DEPTH-1-F1 · the falsification is issue #1328’s comment of 22 Sep, carried into docs/DEPTH-1-NOTES.md on 23 Sep
  • #1398The contamination guard has a hole. A run resumed after the deny list changed can be handed a body an earlier attempt of the same run already fetched, without the guard ever seeing the URL. Narrow — the scorer still refuses to score a run whose store holds a blocked source — but the 16.42 this site publishes rests on that guard.open · BLOCK-1-F24 · crates/kernel/src/retrieval/write.rs
  • #1401main’s CI has been red for 18 consecutive runs since 21 September. Every leg passes but one: macos-15 fails in about seven seconds with no steps recorded, which is a runner or quota problem rather than a test. It gates no number here — the gate checks that run in CI are the Linux leg — but a red main teaches everyone to ignore a red main.open · filed at three runs, 18 as of 2026-09-23 13:54 · gh run list --workflow ci --branch main
  • #1330Four of ten gold-ext runs read the source their own task forbids. Fixed in the code and measured: BLOCK-1 merged on 22 September and enforces the blocklist at the fetch lane by work identity; in the 23 September re-run it refused three fetches, one per mechanism it was built for. The four tasks were re-run without the forbidden source and are four of the ten rows behind the 16.42.the defect is closed in the code; the issue is not · GOLD-EXT-F1 · docs/GOLD-EXT-CLEAN.md

One of the six is worth reading twice. The four tasks re-run without the source their rubric came from scored higher — 16.76 with it, 17.61 without — so the contamination that #1330 found was not an advantage being taken away. It is worth remembering the next time a shortcut looks like one.issues #1326, #1327, #1328, #1330, #1398, #1401, each opened and read on 2026-09-23 · docs/GOLD-EXT-CLEAN.md “What the answer key was worth” · docs/DEPTH-1-NOTES.md

seq 2Notecounted · 2026-09-23

998 findings open, which is the point.

A finding is a defect written down against something already merged. Every package here is read adversarially before it lands and again after, and what that reading turns up is filed rather than fixed on the spot — so the honest total is large, and printing it is cheaper than explaining it away.

laneopen findingsshare
lane:kernel33934 %
lane:eval26426 %
lane:brain14715 %
lane:retrieval10711 %
lane:frontends758 %
lane:governance505 %
lane:ops131 %
lane:doctrine20 %
lane:substrate10 %
total998100 %

Two counting rules disagree and both are printed rather than one being chosen: 998 issues whose title matches [<PACKAGE>-F<n>], and 1,053 carrying the type:finding label, out of 1,114 open issues in all. Most are narrow — a surviving mutant, an untested branch, a stale TESTPLAN row — and none of them is a reason to trust a number here less, because every one is a defect this project found in itself and wrote down. The six above are different only in being visible from outside.gh issue list --state open --limit 2000 over CPUtester5465/research-stack, 2026-09-23, grouped by the lane: label; the finding titles are the adversarial reading's own convention

the other side of this page: what it does today — everything merged and running, with the file each row is read from. What is being built about the six: the roadmap.