research-stack · research.devclusterai.com
MVP · records to 2026-09-23 · every number is from the project's records, listed at the end
worked examplex2-c09-q1run a7a32519ce86

One run, opened all the way down.

A real question from the MVP gate, asked on 2026-09-19. The path it took, the report it wrote, one sentence of that report traced to the exact bytes it rests on, and what it said it could not establish. Ninety-nine seconds; twenty cents.

seq 1Questionas asked

The question — and the one the run decided to answer instead.

Does ignoring robots.txt create legal liability for a research crawler in the US and the EU after the 2024–2025 case law, and where do practitioners draw the safe line? as enrolled · bench/sets/campaigns-mvp.json · campaign x2-c09, “Crawling, robots.txt and archives”
How do US and EU laws differentiate liability for ignoring robots.txt based on the crawler's purpose (commercial vs. non-commercial research), and what technical boundaries define the current practitioner safe line? the run's own rewrite, written in the recharter state and printed at the head of its report
  • runa7a32519ce86… · job x2-c09-q1 · ended done2026-09-19, 10:22–10:23 America/Toronto — first search to last page read
  • depthstandard — a $0.40 ceiling on the run, of which it spent $0.202613bench/sets/gold-8/caps.toml; the reservation is taken before the call, not reconciled after it
  • plan · doctrine2f351b2f373d · 750c8e80cbddcontent hashes of the audit-first plan v1 and doctrine v1 — the two hashes that make this run replayable
  • modelsGemini 3.7 Flash and Gemini 3.1 Pro through OpenRouter, a pair per uniteach recorded as a dated permaslug with a price ref, so the price charged is on the record beside the call

The rewrite is not decoration. recharter is a state with a model call and a cost of its own — $0.038980, the second most expensive state in the run — and the question it writes is the one everything after it is held to.gate ledger: run_head, run_task, state.entered · rsk research result --run x2-c09-q1, head of the report

seq 2Runthe path

Six states, fourteen searches, twenty-four pages, ninety-nine seconds.

run a7a32519ce86 · the states it passed · bar = what each cost 99.0 s · $0.202613 audit $0.035078 1 call 2 assumptions orient $0.017704 1 call 3 searches 8 page reads 5 surprises recharter $0.038980 1 call the rewrite investigate $0.075243 7 calls 11 searches 23 page reads 18 claims synthesise $0.035608 1 call 1 answer · 5 gaps done 99.0 s $0.202613 in all source: the X2 gate ledger — run_task, unit_call, budget_reservation, run_retrieval
whatcountthe row it is read from
states entered6audit → orient → recharter → investigate → synthesise → done — one state.entered row each
model calls11 paid · 0 memounit_call. A memo hit costs no call and no reservation; this run had none, the gate had 5 over 53 runs
searches14search_run — all Brave, one of them returning empty, 105 results between them
pages requested24 distinct31 retrievals over 24 URLs: 7 of the 31 were a page the store already had, handed back with no request and no cost
pages read22the 22 that returned a body — and exactly the 22 sources the report lists
pages refused2onlinelibrary.wiley.com 403 challenge, recovered through web.archive.org by the fallback ladder; iptc.org's PDF answered 200 with unsupported_content_type and was never read
model output rejected1a verifier output refused invariant.A1 — url not fetched in run at /anchors/1/url; the unit ran again and the second attempt was admitted
claims182 accepted · 1 contested · 15 open — the fold's own claim_status rows, not the model's opinion
verdicts3each an adversarial pass carrying its supporting and its disconfirming quotes
wall time99.0 srun.startedrun.finished · docs/GATE-MVP.md, where this run is one line of 53
spend$0.202613budget_reservation — 11 of 11 settled from each response's own usage, in integer micro-USD

The row worth stopping on is the rejected one. A verifier proposed a claim anchored to a page its own task had never fetched. The ledger refused the whole output inside the commit transaction, kept the raw bytes as evidence, and made the unit run again. That second attempt produced the verdict opened below.gate ledger: rejection (stage invariant, code invariant.A1, path /anchors/1/url), unit.rejected and unit.repaired on task b5dc009d… · docs/SPEC-kernel.md, invariant A1

seq 3Argumentthe report

The report it wrote, as it renders.

# Does ignoring robots.txt create legal liability for a research crawler in the US
  and the EU after the 2024-2025 case law, and where do practitioners draw the safe line?

_depth: standard · plan: 2f351b2f373d · doctrine: 750c8e80cbdd · claims: 18 ·
 sources: 22 · calls: 11 · cost: $0.2026 · robots: standard_

## Answer

In the US, ignoring robots.txt to scrape public, unauthenticated data does not
inherently violate the CFAA or breach browsewrap terms (Meta v. Bright Data,
2024). However, it can contribute to DMCA anti-circumvention claims if combined
with technical evasion like IP rotation. In the EU, liability depends entirely
on the type of research. Under the 2025 Kneschke v. LAION ruling, commercial
Text and Data Mining (TDM) must respect machine-readable opt-outs like
robots.txt (Article 4 DSM Directive). Conversely, genuine non-commercial
scientific research is exempt from these opt-outs (Article 3 DSM Directive) and
can legally bypass them. Practitioners draw the safe line by strictly separating
commercial and non-commercial crawlers, adhering to machine-readable opt-outs
for commercial projects, and never bypassing active technical barriers (like
CAPTCHAs or logins) in either jurisdiction.

_Confidence: high confidence._

## Findings

- **In Meta Platforms v. Bright Data (January 2024), the court ruled that
  logged-off public scraping does not breach platform terms of service intended
  for account holders.** - high confidence · dataimpulse.com, cloro.dev
  > Meta v. Bright Data (January 2024) sharpened the line: a court declined to
  > find Bright Data in breach for scraping public, logged-out Facebook and
  > Instagram pages, since the platforms' terms govern account holders, not
  > logged-out visitors.

That is the head of a report of 119 lines. The rest is eleven more findings each with its own quote, the contested claim below, two adversarial verdicts, what it could not establish, next steps, the assumptions the question was resting on, five orientation surprises each marked unverified: no quote from an opened page, and 22 numbered sources. None of it is stored. rsk research result --run x2-c09-q1 derives the whole text from the ledger each time it is asked — which is why replaying a run reproduces its report exactly.the output of rsk research result over the X2 gate ledger, wrapped to fit · docs/SPEC-kernel.md §8, the report interface · docs/GATE-POC.md check 3 (replay: 79 of 79 calls served from the ledger, identical report, $0)

seq 4Fetchone sentence, opened

One sentence of that report, opened.

Following the 2024 Meta v. Bright Data ruling, ignoring robots.txt to scrape publicly accessible, unauthenticated data does not violate the US CFAA or constitute a breach of browsewrap terms. claim d8f8285c… · verdict d2159e30… — the sentence the answer's first clause rests on
one claim · five quotes · each re-read byte for byte at commit verdict: contested “Following the 2024 Meta v. Bright Data ruling, ignoring robots.txt to scrape publicly accessible, unauthenticated data does not violate the US CFAA or constitute a breach of browsewrap terms.” the stored page · byte 0 → end supporting sociavault.com 7,953–8,201 of 33,334 supporting dataimpulse.com 4,570–4,929 of 19,203 supporting cloro.dev 5,825–6,021 of 26,122 disconfirming sociavault.com 8,267–8,660 of 33,334 disconfirming tendem.ai 3,353–3,613 of 10,653 blue supports the claim, red contradicts it — the track is the whole page, the mark is the quote the verifier ruled confirmed; the fold left the claim contested, because two of its own quotes attack it source: arg_anchor rows of verdict d2159e30… — resolved = 1, rung = exact, occurrences = 1, all five the two sociavault.com ranges were re-read from the stored body sha256 0ee390e5… on 2026-09-23
sidepagebytesthe quote, as the ledger stores it
supportingsociavault.com7,953–8,201
of 33,334
“CFAA claim — DISMISSED. The court ruled that scraping publicly available data from Facebook and Instagram does not violate the CFAA. This was consistent with hiQ v. LinkedIn and Van Buren. There's no “unauthorized access” when the data is public.”
supportingdataimpulse.com4,570–4,929
of 19,203
“Meta v. Bright Data (January 2024) sharpened the line: a court declined to find Bright Data in breach for scraping public, logged-out Facebook and Instagram pages, since the platforms’ terms govern account holders, not logged-out visitors. The rule that emerged — logged-out public scraping is defensible; logged-in scraping against accepted terms is not.”
supportingcloro.dev5,825–6,021
of 26,122
“The court held that Meta’s terms govern “your use” of its products — and that Bright Data did not “use” Facebook when it scraped public logged-off pages after terminating its accounts.”
disconfirmingsociavault.com8,267–8,660
of 33,334
“Breach of contract (TOS) — PARTIALLY SURVIVED. The court allowed Meta's breach of contract claim to proceed, but only for the period when Bright Data had an active contractual relationship with Meta. Bright Data had previously been a Meta partner, so they had a direct contract. The court found that the general TOS you “agree to” by browsing a website was a weaker basis for a breach claim.”
disconfirmingtendem.ai3,353–3,613
of 10,653
“Meta sued Bright Data for scraping content from its platforms. The ruling addressed contract-based theories, indicating that scraping content subject to terms of service restrictions may constitute breach of contract even when data appears publicly accessible.”
  • resolved5 of 5, every one rung = exact, occurrences = 1a quote found twice in a page, or not found at all, does not resolve — and an unresolved anchor carries no byte range to print
  • re-readinside the commit transaction, against the stored bodyinvariant A1: an output whose quote is not in the page it cites is never admitted. Not a warning afterwards — a refusal
  • first attemptrefused invariant.A1 — url not fetched in runthe verifier cited a page its own task had not retrieved. $0.010613 paid, nothing admitted, the raw output kept as evidence
  • second attemptadmitted — the five anchors above$0.011012. This one sentence cost $0.021625, of which half bought a refusal
  • the verifier saidconfirmed — three supporting quotes against two contradictingthe model's own label, written into the verdict row and kept there
  • the ledger sayscontested — the claim has two undecided attackersthe fold decides a claim's status, not the model; the report prints the disagreement instead of settling it quietly

This object is the whole argument of the project. Any sentence of any report opens the same way: to its claim, to the quotes under it, to the page each quote came from, to the byte range inside that page, to the verdict, and to what the verdict cost. The two sociavault.com ranges above were decompressed from the body store and re-read while this page was being written; they are the sentences printed beside them, character for character.gate ledger: arg_verdict, arg_anchor, claim_status, budget_reservation, rejection for run a7a32519ce86… · body sha256_text 0ee390e5…, 33,334 bytes of extracted text · docs/SPEC-kernel.md A1 · docs/GATE-POC.md check 4 — 498 anchors, 490 resolved, 0 failing the byte re-read, 0 at a non-ok fetch

seq 5Argumentthe gaps

What it could not establish — in its own words.

  • searched, absent“Searched for explicit cases where non-commercial research organizations bypassed robots.txt and faced litigation specifically testing whether robots.txt is binding under Article 3 DSM Directive / Section 60d UrhG; the available rulings focused on natural language vs. machine-readable opt-outs under Article 4 / Section 44b UrhG.”
  • searched, absent“Searches in the provided text yielded no evidence that breach of browsewrap terms alone satisfies CFAA 'without authorization' thresholds post-Van Buren.”
  • searched, absent“Searching specifically for statutory robots.txt codification in the CFAA yielded nothing; robots.txt remains a machine-readable protocol without independent statutory force under 18 U.S.C. § 1030.”
  • extract too thin“Searched Medium post on AI-oriented robots.txt by Francisco A. Kemeny (file extract provided only generic allow/disallow summary).”
  • extract too thin“Searched IPTC Generative AI Opt-Out Best Practice Recommendations full text beyond snippet (file omitted from retrieved text except snippet).”
  • could not read“Could not read iptc.org — failed (HTTP 200)”the PDF the line above wanted: 200 OK, and a content type the extractor of the day could not open. PDFs are read on main since QUAL-READ, #1203
  • could not read“Could not read onlinelibrary.wiley.com — failed (HTTP 403)”a bot challenge. The fallback ladder recovered the paper through web.archive.org, and it is source 16 of the report

Seven lines, quoted exactly. The report prints nine: two of them restate an earlier line in shorter form, a duplication QUAL-REPORT fixed on 2026-09-21, two days after this run. Nothing else is left out. This is the section no other tool shows you, and it is the reason to believe the eleven findings above it — a run that names what it went looking for and did not find is a run whose silence elsewhere means something.the “What we could not establish” section of rsk research result --run x2-c09-q1, verbatim · docs/PACKAGES.json: QUAL-REPORT (de-duplicated gap lines) and QUAL-READ (#1203, PDFs and query-focused windows), both merged 2026-09-21

next: the same machine at fleet scale, the reference — every part of it that is merged — or try it.