AI crawler policy › domain suffixes

AI crawler policy by domain suffix

Latest sealed date 2026-08-07.

126 suffixes have enough panel domains and parseable robots.txt observations to support a page. Each row keeps the observed denominator next to the two policy counts; an unreachable domain is excluded, never read as permissive.

Suffix means the final DNS label, not the effective top-level domain. A .co.uk site is counted under .uk. The detail pages add a dated series and aggregate verdict changes, so the latest cross-section can be checked against the archive rather than read as a timeless rule.

What the denominator is. On 2026-08-07 we attempted /robots.txt on 100,000 panel domains. 24,130 could not be reached at all (DNS failure, TLS failure, timeout, refused connection) and 14,318 answered with something other than a 200. Both of those groups have no observed AI-crawler policy, and neither is counted as allowing anything — an unreachable site is UNKNOWN, not permissive. Every crawler figure below is out of the 61,552 domains that returned a robots.txt we could parse on that date.
SuffixPanel domainsParsedNames any tracked crawlerBlocks any tracked crawler
.com45,70729,3666,8115,187
.net6,3532,667569491
.org4,4323,243739637
.ru3,6212,186257199
.de2,0001,509450366
.jp1,990842189139
.io1,9611,10115183
.br1,531922163107
.in1,352713138101
.uk1,311889269213
.cn1,1852341510
.fr1,009723212175
.co894568143112
.edu8496967253
.it846584170145
.pl771615186164
.nl609451130109
.au5943939871
.es54834210985
.tv4973586246
.xyz4591503935
.ca4493316351
.me4472825344
.gov4362922522
.id4362704543
.ai4283526836
.app4092474634
.info3922174843
.cz3863117361
.ua3602634233
.eu3461923420
.cloud3409696
.cc3371755049
.tr3262194331
.tw3221022621
.kr3212085445
.ar3151782820
.se3082476652
.vn2992053323
.gr2912184538
.mx2821643826
.online2681363432
.ro2652004934
.ch2601915141
.za2591774337
.hu2551962823
.live2381283530
.pro2371312421
.us230932420
.at2251593730
.be2241564034
.cl2241483019
.site2201012017
.tech2175742
.biz216701212
.fi2161594947
.top202771212
.to2011373938
.pt1841284739
.my1761041512
.dev16984137
.pk1621071613
.ir1598732
.ph156883430
.dk1551264235
.pe154942313
.sk1501162423
.il149104179
.no1441195045
.kz1401021510
.vip136822827
.bg131941310
.link1275276
.by121951211
.club120731210
.gg120942218
.ie119812824
.rs117931210
.nz112793023
.hr103751412
.mobi10366109
.th985476
.blog97781514
.space97371111
.sg92622115
.lt91651515
.one90531410
.az875422
.uz846875
.shop823555
.hk804675
.su803944
.sa784454
.lv765697
.media744533
.xxx736655
.ec723665
.ee675174
.fm6749109
.fun663499
.store663698
.bd653054
.games652766
.ke613686
.ma583176
.ws582987
.ae563564
.st553000
.ng543687
.si543763
.life512866
.is503711
.uy4925105
.world492687
.im463421
.bet442877
.video442722
.am413265
.news41351413
.cat393066
.chat3929114
.ge372733
.ly362777
.md342974
.porn322876
.tube272522

Get told when a domain suffix change their AI crawler rules

Sites publish today's robots.txt and overwrite it — no archive, no history, no notification. We keep a dated daily copy. Leave your email and we'll send you what moved for a domain suffix: who started blocking, who stopped, and on which day.

Email me changes for a domain suffix →

Work email only. No newsletter, no obligation.

Prefer to just buy it? See the sealed-archive monitoring offers, priced, with a dated sample.

Read this before quoting a trend. We hold 55 sealed daily snapshots between 2026-06-09 and 2026-08-07, out of 60 calendar days. 5 calendar day(s) inside that span have no collection run at all: 2026-06-11, 2026-06-12, 2026-06-19, 2026-07-04, 2026-07-15. Consecutive rows in every table below are consecutive SEALED dates, not consecutive days — where a gap falls, that row's diff covers more than 24 hours, and the span is printed with it. We do not interpolate across a day we did not collect.
Why you will not find a robots.txt file here. These pages render mechanical verdicts and aggregate diffs, never a fetched body. Site operators leave contact addresses in robots.txt comments — on the 2026-08-06 snapshot 2,200 of the distinct bodies we fetched contain an email-shaped string — and republishing bodies would republish those. Nothing on this surface identifies an individual site, either: every figure is a count.

Compiled from each site's own published /robots.txt, /ai.txt, /llms.txt and /.well-known/tdmrep.json, fetched daily over a frozen 100,000-domain research panel (source list: Tranco; no Tranco rank is republished here). Verdicts are mechanical readings of what those files say under RFC 9309 grouping with exact product-token matching — they are not legal advice, not a statement about any site's intentions, and not a claim about what any crawler actually did. A robots.txt is a request, not an access control. Figures are provided for informational purposes only and carry no warranty of accuracy or completeness.

AI crawler policy home · What changed · How this is collected · US Tech Automations.