# ENCLOSURE DOCKET — Sample Evidence Pack v0

**v0 — PUBLISHED 2026-07-10 under dated operator approval. Sealed and OpenTimestamps-anchored; every number is re-derivable from source via the product's `verify.py`.**

| | |
|---|---|
| Subject domain | **flickr.com** (Tranco rank 173 in universe `tranco_L8QG4_top100000.txt:8778484edc1487aa`) |
| Event | robots.txt AI-crawler access policy change |
| Bracketed between | snapshot **2026-06-29** (10:02:39 UTC) and snapshot **2026-06-30** (10:05:35 UTC) |
| Direction | **Selective opening** — five AI user/search agents moved from fully blocked to allowed |
| Panel | Tranco top-100k domains × 4 policy surfaces × 27 daily snapshots, 2026-06-09 → 2026-07-09 |
| Pack generated | 2026-07-10 by USTA closing-web corpus (read-only extraction) |
| Machine claims | `data.json` (this directory), all re-derivable via `verify.py` — 71/71 PASS at build time |

---

## 1. Finding (plain English)

Between the daily snapshots of **2026-06-29 10:02 UTC** and **2026-06-30 10:05 UTC**, `flickr.com`
changed its `robots.txt` from content hash `1152a640…` to `7f0af7ac…`. The change is exactly
**5 added lines, 0 removed lines**. All five additions are AI-agent `User-agent` tokens inserted
into Flickr's *named-crawler allow group*:

- `ChatGPT-User` (OpenAI, acts on behalf of ChatGPT users)
- `Claude-SearchBot` (Anthropic, search indexing)
- `Claude-User` (Anthropic, acts on behalf of Claude users)
- `OAI-SearchBot` (OpenAI, search indexing)
- `PerplexityBot` (Perplexity, search indexing)

Why this matters: the *final `User-agent` group* in Flickr's robots.txt is
`User-agent: * / Disallow: /` — a **default-deny** for every crawler not explicitly named. (The
only lines after that group, in both versions, are five `Sitemap:` lines, which carry no access
rules — machine-checked as `tail_structure_before/after`.) Before 2026-06-30, none of these five agents was named, so
all five were fully blocked. After the change, they are permitted site-wide except ~19 path-scoped
`Disallow` rules shared by the named group (search pages, lightboxes, OAuth, several member pages).

Equally load-bearing, and verified in both versions: the **training crawlers were never named** —
`GPTBot`, `ClaudeBot`, `CCBot`, `Google-Extended`, `Bytespider`, `anthropic-ai`,
`meta-externalagent` appear in *neither* version and therefore remain blocked by default-deny
throughout the window. Flickr opened the door to AI *answer/search* agents while keeping AI
*training* crawlers shut out — a selective policy change on a major photo-hosting platform whose
library includes large volumes of copyrighted photographs. (What this pack proves is the published
policy text and its timing. Intent and motive are not observable in robots.txt — see caveat 5 —
and we cite no figure for the size of Flickr's photo library, which we did not source.)

Companion surfaces did not move: `/ai.txt`, `/llms.txt`, and `/.well-known/tdmrep.json` returned
**404 on all 27 snapshots**; robots.txt itself returned **HTTP 200 on all 27 snapshots** (no
403/429 gating observed).

Corroboration (informational): a live fetch of `https://www.flickr.com/robots.txt` on 2026-07-10
hashed to `7f0af7ac…` — byte-identical to the stored "after" document.

## 2. Dated hash timeline — flickr.com, resource `robots`

27 snapshots, all HTTP 200. Gaps in the daily panel: 2026-06-11, 06-12, 06-19, 07-04 (no
collection run those days). Full 64-hex values are in `data.json` (`flickr_timeline`).

| snapshot_date | status | content_sha256 (12) | row_sha256 (12) | collected_at (UTC) |
|---|---|---|---|---|
| 2026-06-09 | 200 | `1152a6402b72` | `f3ab7ed786f7` | 19:43:27 |
| 2026-06-10 | 200 | `1152a6402b72` | `80552408eae3` | 10:08:06 |
| 2026-06-13 | 200 | `1152a6402b72` | `4fc8d66f31c6` | 10:07:46 |
| 2026-06-14 | 200 | `1152a6402b72` | `b01932ad3ce7` | 10:03:50 |
| 2026-06-15 | 200 | `1152a6402b72` | `d163470ff5e2` | 10:09:53 |
| 2026-06-16 | 200 | `1152a6402b72` | `d5b19db60083` | 10:00:29 |
| 2026-06-17 | 200 | `1152a6402b72` | `7a83d4dba5be` | 10:02:09 |
| 2026-06-18 | 200 | `1152a6402b72` | `8ae775fd23c6` | 10:09:36 |
| 2026-06-20 | 200 | `1152a6402b72` | `4f97b16a7fb0` | 18:54:38 |
| 2026-06-21 | 200 | `1152a6402b72` | `0b5f76e080f3` | 10:03:50 |
| 2026-06-22 | 200 | `1152a6402b72` | `278d716df521` | 10:08:46 |
| 2026-06-23 | 200 | `1152a6402b72` | `1ce66c656567` | 10:05:50 |
| 2026-06-24 | 200 | `1152a6402b72` | `bb27b6eda8c5` | 10:09:50 |
| 2026-06-25 | 200 | `1152a6402b72` | `27299640f40f` | 10:09:50 |
| 2026-06-26 | 200 | `1152a6402b72` | `e2836362135f` | 10:00:42 |
| 2026-06-27 | 200 | `1152a6402b72` | `eeedd05036c8` | 10:05:37 |
| 2026-06-28 | 200 | `1152a6402b72` | `8df1e2b8087d` | 10:09:52 |
| **2026-06-29** | 200 | **`1152a6402b72`** ← last "before" | `af21aba3c6de` | 10:02:39 |
| **2026-06-30** | 200 | **`7f0af7ac6cba`** ← first "after" | `d3694d377a6a` | 10:05:35 |
| 2026-07-01 | 200 | `7f0af7ac6cba` | `54627d9b3e9e` | 10:06:35 |
| 2026-07-02 | 200 | `7f0af7ac6cba` | `442d614b3a53` | 10:05:29 |
| 2026-07-03 | 200 | `7f0af7ac6cba` | `d9307df6106f` | 15:39:40 |
| 2026-07-05 | 200 | `7f0af7ac6cba` | `47ae42fc1003` | 20:53:29 |
| 2026-07-06 | 200 | `7f0af7ac6cba` | `03e396ecdfa8` | 10:01:57 |
| 2026-07-07 | 200 | `7f0af7ac6cba` | `8d97b1a4f418` | 10:03:07 |
| 2026-07-08 | 200 | `7f0af7ac6cba` | `d21c2895bfaa` | 10:04:20 |
| 2026-07-09 | 200 | `7f0af7ac6cba` | `2771f5410bfb` | 10:06:42 |

Exactly **one** hash transition occurs in the window (machine-checked:
`flickr_robots_hash_transition_count = 1`).

Full hashes of the two documents:

- BEFORE `content_sha256` = `1152a6402b728a8bb89ace366a112ede4cbd1be82ecac8734eea9cdeb20fc617` (2,000 bytes)
- AFTER  `content_sha256` = `7f0af7ac6cba1389d7f1271821f9c4330b258c1e387803936a00442d19aae402` (2,130 bytes)

## 3. Before / after content and diff

Full documents are preserved byte-exact in this pack as
`before_robots_2026-06-29.txt` and `after_robots_2026-06-30.txt`
(robots.txt is public content; sha256 of each file equals the content hash above).

Structure of both versions (excerpt, "after" shown):

```
User-agent: Adsbot-Google
User-agent: Alexabot
...
User-agent: BingPreview
User-agent: ChatGPT-User        <-- added 2026-06-30
User-agent: Claude-SearchBot    <-- added 2026-06-30
User-agent: Claude-User         <-- added 2026-06-30
User-agent: coccocbot
...
User-agent: NAVER
User-agent: OAI-SearchBot       <-- added 2026-06-30
User-agent: PerplexityBot       <-- added 2026-06-30
User-agent: PetalBot
...
User-agent: ZoomBot
Disallow: /gp/
Disallow: /report_abuse.gne
...  (19 path-scoped Disallow rules for the named group)

User-agent: magpie-crawler
Disallow: /

User-agent: *
Disallow: /          <-- default-deny: everything not named above is blocked
```

Unified diff (complete — this is the entire change; also stored as `robots_diff.patch`):

```diff
--- robots.txt@2026-06-29 (1152a6402b72)
+++ robots.txt@2026-06-30 (7f0af7ac6cba)
@@ -7,6 +7,9 @@
 User-agent: bitlybot
 User-agent: bingbot
 User-agent: BingPreview
+User-agent: ChatGPT-User
+User-agent: Claude-SearchBot
+User-agent: Claude-User
 User-agent: coccocbot
 User-agent: Discordbot
 User-agent: DuckDuckBot
@@ -24,6 +27,8 @@
 User-agent: Mediapartners-Google
 User-agent: msnbot
 User-agent: NAVER
+User-agent: OAI-SearchBot
+User-agent: PerplexityBot
 User-agent: PetalBot
 User-agent: Pingdom
 User-agent: Pinterest
```

## 4. Chain of custody

The corpus commits to observations in three internal layers plus one (limited) external layer.
Every internal link below was re-derived from raw bytes in this run and is re-derivable by
`verify.py`.

**Link 1 — content hash.** Response bodies are stored zlib-compressed, content-addressed in
`blobs(content_sha256, content_gz, …)`. We decompressed both documents and recomputed sha256 over
the raw bytes; both equal the stored `content_sha256` values. ✔

**Link 2 — row hash.** Every snapshot row carries
`row_sha256 = sha256(canonical_json({domain, snapshot_date, resource, status_code, content_sha256,
headers_json, fetch_error}))` (canonical JSON: sorted keys, `,`/`:` separators, ASCII — per
`closing_web/store.py::build_row`). Recomputed for the two key rows:

- 2026-06-29 row: `af21aba3c6de47de943aa7daad4f66446b4ae5a6ff61a6d947c2a8dd4e950e9d` ✔
- 2026-06-30 row: `d3694d377a6a18e67999b4d260e34a4adfd1b415bbf202604effb6225ef7d0cb` ✔

The stored allowlisted headers for both rows are `{"content-type":"text/plain;charset=UTF-8"}`;
`fetch_error` is NULL. The row hash binds the content hash to the date, so the timeline in §2
cannot be altered without breaking every row hash.

**Link 3 — per-run batch commitment.** Each collection *run* writes
`batch_sha256 = sha256(canonical_json(sorted([row_sha256 …])))` over the rows **that run
inserted** into `collection_runs`. Batches are **per-run, not per-day**: the window contains
**29 runs across 27 snapshot days** (2026-06-09 and 2026-06-13 each had two runs). All 29 stored
`batch_sha256` values re-derive byte-exact from their run's row hashes (rows carry
`collected_at` equal to the run's `started_at`) ✔. The two key days are each a single
full-panel run:

- Run `2026-06-29:2026-06-29T10:02:39.718325+00:00` — 400,000 rows (the day's entire row set),
  `batch_sha256 = acdb51b2a95418203ad4b5165e6f50f28f813a552ff6e93c0d41350a00ba122e` ✔
  (contains the 06-29 key row hash ✔)
- Run `2026-06-30:2026-06-30T10:05:35.698059+00:00` — 400,000 rows (the day's entire row set),
  `batch_sha256 = 1a1bad3abd226204c18274c738720c40e4616ff5e0e742dc4e7d88fde163d744` ✔
  (contains the 06-30 key row hash ✔)

So the two rows that bracket the flip are each provably members of their run's 400,000-row
batch: tampering with either row changes that batch hash. **This membership claim is scoped to
the days verified above — it does NOT hold corpus-wide**, because of the following disclosed gap.

**Disclosed layer-3 gap (2026-06-13) and the 06-09 split.** The union of all 29 run batches
commits to **10,410,000 of the 10,800,000 rows** in the window
(`sum_rows_inserted_all_runs`); the 390,000-row shortfall is entirely on 2026-06-13
(machine-checked day by day, `days_with_uncommitted_rows`):

- **2026-06-09** — two runs (200 rows + 399,800 rows); both batches re-derive ✔ and flickr's
  06-09 timeline row is inside the 399,800-row batch ✔ (`flickr_row_2026-06-09_covering_run_id`).
  A split run, but full coverage.
- **2026-06-13** — the day holds 400,000 rows, but its two runs commit to only **0 rows** (the
  stored batch hash is byte-exactly sha256 of an empty list ✔) and **10,000 rows** ✔. The
  remaining **390,000 rows — including flickr's 2026-06-13 timeline row — carry no layer-3
  batch commitment** (`flickr_row_2026-06-13_committed_by_any_run = false`). For those rows,
  links 1–2 (content and row hashes) still hold, but tampering with one of them would not break
  any batch hash. This does not touch the flip evidence (06-29 and 06-30 are fully committed),
  and flickr's 06-13 content hash equals the 17 other batch-committed "before" days. See
  caveat 9.

**Link 4 — external commitment (what it does and does NOT prove).** Being exact here is the
product:

- A monthly derived seal exists: `closingweb-tranco-100k-2026-07`, declared
  `sha256 = 007e5b10ad213ec045505c1f7b0a851f932f80f61bf682bd0d3ff74a9ae5f130`, last sealed
  2026-07-09 (file: `~/code/USTA/seo/data/research/snapshots/closingweb-tranco-100k-2026-07/snapshot.json`).
  **Limitation:** this seal hashes an *aggregate statistics snapshot* (counts, percentages)
  derived from the corpus. It does **not** commit to per-row hashes, per-domain hashes, or batch
  hashes. Per-row inclusion of the flickr observations in this seal is **not provable**.
- An OpenTimestamps-anchored manifest exists:
  `~/Claude CLI/provenance-anchor/manifests/priority_manifest_2026-07-09.json`
  (file sha256 `1d96bebdfe57a30f89fd711f3cb40a7331abc230874856ffcd9f0002c745c271`, 405 entries,
  merkle root `e2cd10ac9421bb03a5fcda75a78a02e1a3d69a6e4cb706dc07f993d66b434571` — merkle root
  and internal `manifest_content_sha256` re-derived and confirmed in this run), with a Bitcoin
  `.ots` attestation beside it. **Limitation:** we machine-verified that this manifest contains
  **zero** entries referencing the closing-web corpus (its 405 entries cover other USTA artifact
  families). Therefore the Bitcoin timestamp proves **nothing** about the flickr rows. As of this
  pack, no closing-web row, batch, or seal is externally anchored. Bitcoin-inclusion of the `.ots`
  itself was not independently verified in this run (requires the `opentimestamps` client;
  marked unverifiable here).

**Net custody claim, stated honestly:** the flickr flip is protected by a consistent three-layer
*internal* hash chain in an append-only store maintained by an automated daily collector (layers
1–2 complete across the window; layer-3 batches cover 10.41M of 10.8M rows, with the 06-13 gap
disclosed above), and the "after" document is independently corroborated by the live web today
(and by anyone who fetches `flickr.com/robots.txt` while the policy holds). Third-party web
archives are a *plausible additional* corroboration path, but no archive capture was fetched or
hashed for this pack — it is a lead, not evidence in hand. The flip is **not yet** protected by
an external cryptographic timestamp. Adding closing-web `batch_sha256` values to the
OTS-anchored manifest would close this gap prospectively and is the obvious v1 improvement.

## 5. Market note (addressable inventory)

In this single 31-day window, **14,620** of the Tranco top-100k domains served at least two
distinct robots.txt contents — i.e., at least one observed robots.txt change — out of **63,544**
domains that served fetchable robots.txt content at least once (23.0%). Each such change is a
candidate docket entry; the subset that touches AI-crawler directives is the premium inventory
for litigators, AI-lab compliance teams, and publishers who need to prove exactly when a policy
stood where. A spot-scan of the top 800 changed domains (by Tranco rank), using the explicit
27-token AI-crawler list and hit rule disclosed in `data.json` (`spot_scan_ai_token_list`,
`spot_scan_rule`) and re-derived by `verify.py`, found **69** whose AI-token set changed within
the window — including forbes.com, hulu.com, spiegel.de, khanacademy.org, axios.com, and
shutterstock.com, each machine-confirmed (`spot_scan_examples_all_hit`).

## 6. Methodology

1. **Corpus.** The USTA "closing-web" clock fetches 4 policy surfaces (`/robots.txt`, `/ai.txt`,
   `/llms.txt`, `/.well-known/tdmrep.json`) for a frozen Tranco top-100k universe daily
   (~10:00 UTC), sealing rows append-only (first write per (domain, date, resource) wins) into
   SQLite with the 3-layer hash chain described in §4. Methodology id: `closing-web-v1`.
2. **Candidate detection (this pack).** Grouped `resource='robots'` rows per domain, counted
   distinct non-NULL `content_sha256` (window ≤ 2026-07-09); shortlisted the 14,620 domains with
   ≥2; ranked by Tranco position; decompressed all version blobs for the top 800 and flagged
   domains where the *set of AI-bot tokens* (27-token list in `data.json`) was not constant
   across their date-ordered versions (69 hits, re-derived by `verify.py`); selected
   flickr.com from three finalists (flickr, forbes, hulu) for: single clean transition, minimal
   all-AI diff, and copyright salience. Forbes was rejected because its two versions oscillate
   (A/B or CDN-edge inconsistency) — documented as an anti-example, not an event date.
3. **Verification.** Every number in `data.json` re-derived from the raw store by `verify.py`
   (stdlib-only, read-only). 71/71 claims PASS.

## 7. Data sources

| Source | Path | Rows used |
|---|---|---|
| Closing-web corpus (SQLite, read-only) | `~/Claude CLI/clocks/closing_web/data/closing_web.db` | `policy_snapshots`: 10,800,000 rows ≤ 2026-07-09 across **27 snapshot days**; flickr robots: 27 rows; `blobs`: 2 key documents + top-800 spot-scan versions; `collection_runs`: **29 runs** (all 29 batch hashes re-derived; union covers 10,410,000 rows — see §4 Link 3 gap) |
| Frozen universe | `~/Claude CLI/clocks/closing_web/universe/tranco_L8QG4_top100000.txt` | 100,000 (flickr rank 173) |
| Monthly derived seal | `~/code/USTA/seo/data/research/snapshots/closingweb-tranco-100k-2026-07/snapshot.json` | 1 declared sha256 |
| Anchored manifest + OTS | `~/Claude CLI/provenance-anchor/manifests/priority_manifest_2026-07-09.json` (+ `.ots`) | 405 entries (0 closing-web) |
| Hash-chain spec | `~/Claude CLI/clocks/closing_web/closing_web/store.py` | build_row / seal logic |

## 8. Reproduction

```bash
cd "~/Claude CLI/exhaust-products/enclosure_docket"
python3 verify.py            # re-derives all 71 claims from the sources above; exit 0 = all PASS
                             # (~30 s: re-hashes all 29 run batches + top-800 spot-scan)
python3 verify.py --live     # additionally re-fetches flickr.com/robots.txt (informational)
python3 build_data.py        # regenerates data.json + evidence files from scratch
```

`verify.py` is stdlib-only, opens the database with `file:...?mode=ro`, and prints a per-claim
PASS/FAIL table. The window is frozen at ≤ 2026-07-09 so results are stable as the live corpus
grows. `data.json` field `_artifact_sha256` holds sha256 of this markdown file (computed over the
file without that field, i.e., the hash is of the `.md` bytes only).

## 9. Caveats / limitations (read before relying on this)

1. **n = 1 sample.** This pack documents one domain's one change. It demonstrates the evidence
   format; it is not a corpus product yet.
2. **Short panel, coarse clock.** 31 days, one probe per day (~10:00 UTC), 4 gap days (06-11,
   06-12, 06-19, 07-04). The flip is bracketed to a ~24h window (2026-06-29 10:02 → 2026-06-30
   10:05 UTC), not timed to the minute. The change could have occurred anytime in that window.
3. **No external timestamp on this data (yet).** The Bitcoin-anchored manifest does not cover
   closing-web rows (§4, verified zero entries). All custody below the live-web corroboration is
   internal to one operator's append-only store. A court-grade opponent could argue the store was
   fabricated wholesale; the rebuttals today are (a) internal consistency across 10.8M rows at
   layers 1–2 — qualified by the layer-3 gap in caveat 9 — and (b) the live web currently
   matching the "after" hash. Third-party archives (e.g., a Wayback capture of flickr's
   robots.txt inside the window) would be a third rebuttal, but **no archive capture was fetched,
   hashed, or checked for this pack** — that is prospective work, not an existing rebuttal.
   Prospective OTS anchoring of per-run `batch_sha256` values is the fix.
4. **Single vantage point.** One fetcher, one network location. CDNs can serve different
   robots.txt per region/edge (Forbes visibly did). Flickr's flip shows a clean single transition,
   consistent with a real origin change, but a multi-vantage panel would be stronger.
5. **robots.txt is a request, not an ACL.** This pack proves what policy Flickr *published*
   and when — not whether any crawler honored it, and not Flickr's contractual arrangements
   (the opening to OpenAI/Anthropic/Perplexity agents may reflect a private deal; we have no
   visibility into that).
6. **Interpretation of "AI-crawler tokens" is list-based.** The spot-scan uses the exact
   27-token list and hit rule disclosed in `data.json` and covers only the top 800 of 14,620
   changed domains. The 69-hit count in §5 is exact for that list/rule over the top-800, but a
   floor for the corpus: other domains were not scanned and tokens outside the list are not
   counted.
7. **Discrepancies vs. planning hints.** Orientation hints described "~10.4M rows / ~27 snapshots
   2026-06-09→07-09" and gzipped blobs. Measured: 10,800,000 rows in-window (10.9M+ including the
   in-flight 07-10 run) and **zlib**-compressed blobs (not gzip). Measured values govern.
8. **Live-web corroboration decays.** The 2026-07-10 live match is a point-in-time observation;
   flickr may change robots.txt again at any moment.
9. **Layer-3 batch coverage has a one-day gap.** Batch commitments are per-run, and on
   2026-06-13 the two recorded runs commit to only 10,000 of the day's 400,000 rows — the other
   390,000 rows (~3.6% of the corpus), **including flickr's 06-13 timeline row**, are covered by
   content and row hashes (layers 1–2) but by **no batch hash**. Tampering with one of those rows
   would not break any batch commitment. The flip-bracketing days (06-29, 06-30) are fully
   committed, and flickr's 06-13 content hash equals the surrounding batch-committed days, so the
   finding stands; but any corpus-wide custody statement must carry this exception (all figures
   machine-checked: `days_with_uncommitted_rows`, `uncommitted_rows_2026-06-13`,
   `flickr_row_2026-06-13_committed_by_any_run`).

---

*Prepared read-only from the USTA closing-web corpus. No existing database, script, service, or
timer was modified. All files in this pack are new writes under
`~/Claude CLI/exhaust-products/enclosure_docket/`.*
