AI crawler policy › methodology
Methodology census-v1 · read-model built 2026-08-08T21:01:56Z.
A frozen 100,000-domain research list (tranco_L8QG4_top100000.txt:8778484edc1487aa), sourced from the Tranco top-100k. The list is frozen so that a change in a count is a change in the web and not a change in who we asked. We cite Tranco as the panel source and publish only our own dated observations over it; no Tranco rank appears anywhere on this surface. The panel is not a random sample of the web — it is a popularity-ranked list, so these figures describe large sites and must not be read as "x% of websites".
Four files per domain: /robots.txt, /ai.txt, /llms.txt, /.well-known/tdmrep.json. Response bodies are stored content-addressed by SHA-256 and compressed; identical bytes across days and across domains are stored once, which is what makes a daily archive of this size tractable. Row integrity is verified by a per-row hash at seal time — the archive can prove that the copy we hold for a given day is the copy we fetched that day.
| File | Rows sealed | Status 200 | Genuine file | HTML shell (soft 200) | Empty | Transport error |
|---|---|---|---|---|---|---|
/ai.txt | 100,000 | 12,022 | 427 | 11,338 | 257 | 24,370 |
/llms.txt | 100,000 | 19,051 | 8,155 | 10,588 | 308 | 24,327 |
/robots.txt | 100,000 | 62,134 | 62,134 | 0 | 0 | 24,130 |
/.well-known/tdmrep.json | 100,000 | 10,850 | 178 | 10,418 | 254 | 24,357 |
Many sites answer any unknown path with their HTML application shell and a 200. Counted on the sealed 2026-06-10 snapshot: 17,934 domains returned 200 for /llms.txt, but only 6,841 of those bodies were a genuine text file — 10,786 were HTML shells and 307 were empty. Reporting the 200 count as adoption would have overstated it by 2.6x. Every adoption figure on this surface uses the genuine-only numerator, and the soft-200 and empty counts are printed beside it so you can check the arithmetic — genuine + shell + empty equals the 200 count on every row of that table, by construction. A 200 that carried no body at all is counted as empty rather than dropped: dropping it would leave a decomposition that does not add up to its own total, which is how a denominator quietly stops being one.
On 2026-08-07, 97,184 of the day's fetches across all four files ended in a transport failure rather than an HTTP response. Those domains have no observed policy on that date. They are excluded from every numerator and every denominator on this surface — never counted as "allows everything", and never reported as the site being down, because a fetch we could not complete is a fact about our fetch.
| Failure class | Fetches | Share of failures |
|---|---|---|
dns_error | 55,420 | 57.0% |
timeout | 18,505 | 19.0% |
ssl_error | 16,334 | 16.8% |
connection_refused | 3,711 | 3.8% |
oserror | 1,960 | 2.0% |
connection_reset | 1,189 | 1.2% |
incompleteread | 31 | 0.0% |
invalidurl | 11 | 0.0% |
unicodeerror | 9 | 0.0% |
unicodeencodeerror | 7 | 0.0% |
url_error | 4 | 0.0% |
badstatusline | 3 | 0.0% |
Robots.txt is parsed under RFC 9309 grouping: consecutive User-agent lines open a group and the rules that follow apply to every agent named in it; multiple groups for the same token are merged. A crawler is governed by the group whose user-agent value equals its product token case-insensitively, otherwise by the * group. There is no substring matching and no fuzzy matching anywhere: a group for GPTBot-Mirror is a different token and does not count as a rule about GPTBot.
rows_inserted log does not match the number of rows we can read for that date. Where they disagree we report the rows — a run log is a claim about a fetch, and the rows are the fetch. Dates affected: 2026-06-13. "Partial pass" above therefore means fewer sealed rows than a full day, measured against the same rows every figure on this surface is computed over.We also do not publish per-domain verdicts on this surface. Every page here is an aggregate. A mechanical verdict is a reading of a file; attached to a named site it starts to read as a claim about that company's intentions, which we did not observe and are not entitled to assert. And we do not publish what any crawler actually did — robots.txt is a request, and compliance is not something this archive can see.
Want this cut a different way for the sites you care about?
Tell us the domains, the crawler operators or the date range you care about and we'll pull that slice out of the sealed archive and send it over.
Send me a custom crawler-policy view →Work email only. No newsletter, no obligation.
Prefer to just buy it? See the sealed-archive monitoring offers, priced, with a dated sample.
Compiled from each site's own published /robots.txt, /ai.txt, /llms.txt and /.well-known/tdmrep.json, fetched daily over a frozen 100,000-domain research panel (source list: Tranco; no Tranco rank is republished here). Verdicts are mechanical readings of what those files say under RFC 9309 grouping with exact product-token matching — they are not legal advice, not a statement about any site's intentions, and not a claim about what any crawler actually did. A robots.txt is a request, not an access control. Figures are provided for informational purposes only and carry no warranty of accuracy or completeness.