Six defects a reader of this site can see.
- #1326A run whose writer cannot get its article past the citation gate ends
gap, and the finished research is discarded — 3 of 10 runs in the plan-v2 external pass, each with its synthesis already on the record. Being fixed inside EXTRACT-1: a refused article should leave the rundonewith its report and the refusal named in its meta line.open · NARR-1-F8 · docs/DEPTH-1-NOTES.md, the live pass - #1327The report pads. About 130 of one report’s 260 lines were a list of failed fetches, Findings repeated Sources verbatim, and “what is not established” was printed twice. Presentation scores 27.04 on the clean ten, and this is part of why.open · NARR-1-F9 · said unprompted by the blind pairwise readers
- #1328Analysis fell 47 % — 31.21 to 16.42 over the same eight tasks —
when the cover loop doubled what each run read. Our first explanation for it measured false: we said the synthesist
was covering more claims under a fixed output budget, and the ledger says that budget never bound — 39 of 39 calls ended
stop, the synthesist using at most 3,670 of its 8,000 tokens. The article is claim-starved instead, and 19 of 21 of the mechanism terms the judge missed were sitting in pages the run had already stored.open · DEPTH-1-F1 · the falsification is issue #1328’s comment of 22 Sep, carried into docs/DEPTH-1-NOTES.md on 23 Sep - #1398The contamination guard has a hole. A run resumed after the deny list changed can be handed a body an earlier attempt of the same run already fetched, without the guard ever seeing the URL. Narrow — the scorer still refuses to score a run whose store holds a blocked source — but the 16.42 this site publishes rests on that guard.open · BLOCK-1-F24 · crates/kernel/src/retrieval/write.rs
- #1401main’s CI has been red for 18 consecutive runs since 21
September. Every leg passes but one:
macos-15fails in about seven seconds with no steps recorded, which is a runner or quota problem rather than a test. It gates no number here — the gate checks that run in CI are the Linux leg — but a red main teaches everyone to ignore a red main.open · filed at three runs, 18 as of 2026-09-23 13:54 · gh run list --workflow ci --branch main - #1330Four of ten gold-ext runs read the source their own task forbids. Fixed in the code and measured: BLOCK-1 merged on 22 September and enforces the blocklist at the fetch lane by work identity; in the 23 September re-run it refused three fetches, one per mechanism it was built for. The four tasks were re-run without the forbidden source and are four of the ten rows behind the 16.42.the defect is closed in the code; the issue is not · GOLD-EXT-F1 · docs/GOLD-EXT-CLEAN.md
One of the six is worth reading twice. The four tasks re-run without the source their rubric came from scored higher — 16.76 with it, 17.61 without — so the contamination that #1330 found was not an advantage being taken away. It is worth remembering the next time a shortcut looks like one.issues #1326, #1327, #1328, #1330, #1398, #1401, each opened and read on 2026-09-23 · docs/GOLD-EXT-CLEAN.md “What the answer key was worth” · docs/DEPTH-1-NOTES.md