“A comprehensive survey, in the shape I was given.” — not yet.
A required shape: named tables, fixed sections, a checklist of entities that must all be covered. If that is your case today, do not ask for access. Here is why, in the only terms that should persuade you — a measurement someone else wrote the rubric for.
Ten DeepResearch Bench II tasks, 672 expert-written rubric items, scored under gold-ext/v3 with no run having read the source its own rubric was derived from — the fetch lane refuses those by work identity, and three fetches were refused live during the pass.
| dimension | measured, the clean ten | best published | what it means for this case |
| overall | 16.42 | 45.40 | the whole rubric, 672 items over ten tasks |
| information recall | 12.89 | 39.98 | how much of what was asked for is in the report at all — a survey is mostly this |
| analysis | 27.16 | 49.85 | does it explain and weigh, or only list |
| presentation | 27.04 | 89.16 | the named tables and fixed sections a required shape asks for — the widest gap of the four |
The right-hand column is a different judge. The published table (arXiv:2601.08536v3) was scored by Gemini-2.5-Pro, and its repository later moved to GPT-5.5; ours is Gemini-3.7-flash under our own prompts. The dimensions are comparable in shape, not in calibration, so read the column as a direction and not as a scoreboard. Our own before-and-after comparisons are sound because both sides are scored the same way.docs/GOLD-EXT-CLEAN.md — the ten tasks, the per-task table, the leaderboard row and why it is a different judge
What would change our mind
Not more reading. Asked whether any single passage carried a complete rubric fact, real items scored 0.48–0.64 against fabricated controls at 0.04: a study’s name sits in one place, its country, design and sample size in another. The constraint is assembly, not reading. EXTRACT-1 (#1404) takes statements off the page; ASSEMBLE-1 (#1405) joins them into records.
When a re-run of the clean ten moves information recall off 12.89, this page changes and says so. Until then it says this.docs/JEV-MEASUREMENT.md — the complete-or-partial split, 680 windows, $0.0643 · issues #1404, #1405 · the roadmap