Which of the top 100,000 sites block AI crawlers?

Counted from sealed daily snapshots. Latest: 2026-08-07.

On 2026-08-07 we read /robots.txt from 100,000 domains in the panel. 61,552 returned a file we could parse. Of those, 13,739 (22.3%) name at least one of the 30 AI crawler product tokens we track, and 10,802 (17.5%) disallow at least one of them from the whole site. A further 2,216 block every crawler via the wildcard group, which is not a decision about AI in particular and is counted separately. 24,130 panel domains could not be reached at all and have no observed policy.
What the denominator is. On 2026-08-07 we attempted /robots.txt on 100,000 panel domains. 24,130 could not be reached at all (DNS failure, TLS failure, timeout, refused connection) and 14,318 answered with something other than a 200. Both of those groups have no observed AI-crawler policy, and neither is counted as allowing anything — an unreachable site is UNKNOWN, not permissive. Every crawler figure below is out of the 61,552 domains that returned a robots.txt we could parse on that date.

Every crawler token we track, on the latest sealed date

CrawlerOperatorNames itBlocks outrightPartly restrictsNames & allows% of parsed
GPTBotOpenAI11,0318,5331,1761,32213.9%
ClaudeBotAnthropic9,7577,5819271,24912.3%
CCBotCommon Crawl9,4948,34855459213.6%
Google-ExtendedGoogle8,9277,0408051,08211.4%
BytespiderByteDance8,5587,87939728212.8%
AmazonbotAmazon7,9956,94354051211.3%
meta-externalagentMeta7,7846,74354349811.0%
Applebot-ExtendedApple7,5366,61147545010.7%
PerplexityBotPerplexity4,6082,3609261,3223.8%
ChatGPT-UserOpenAI4,5822,4049781,2003.9%
anthropic-aiAnthropic3,7282,7054965274.4%
OAI-SearchBotOpenAI3,2601,2968931,0712.1%
cohere-aiCohere2,9072,3043502533.7%
Claude-WebAnthropic2,8552,1503133923.5%
FacebookBotMeta2,5931,9533143263.2%
omgiliWebz.io2,5092,303157493.7%
DiffbotDiffbot2,4192,1551461183.5%
omgilibotWebz.io2,2872,088150493.4%
ImagesiftBotImageSift1,8981,76287492.9%
TimpibotTimpi1,6481,51291452.5%
Perplexity-UserPerplexity1,5947803604541.3%
Claude-SearchBotAnthropic1,5837993834011.3%
Claude-UserAnthropic1,4777393443941.2%
AI2BotAi21,3471,129158601.8%
DuckAssistBotDuckDuckGo1,3479042711721.5%
meta-externalfetcherMeta1,2889601751531.6%
cohere-training-data-crawlerCohere1,046910108281.5%
Google-CloudVertexBotGoogle90273693731.2%
MistralAI-UserMistral8476131281061.0%
PanguBotHuawei81877231151.3%

Blocks outright — the file has a group for that exact product token containing Disallow: / with no overriding allow. Partly restricts — a group exists with at least one non-empty Disallow, but not the whole site. Names & allows — a group exists with no effective Disallow. Everything else is not mentioned: the wildcard * group governs it, and wildcard effects are counted separately rather than folded in, because a site that blocks everything has not made a decision about any AI crawler in particular. Matching is exact and case-insensitive per RFC 9309 — never substring.

All 30 tracked tokens are named by at least 250 parsed domains and have their own page. A token below that floor would still appear in this table, with its count, and simply not be linked.

Get told when the sites you care about change their AI crawler rules

Sites publish today's robots.txt and overwrite it — no archive, no history, no notification. We keep a dated daily copy. Leave your email and we'll send you what moved for the sites you care about: who started blocking, who stopped, and on which day.

Email me AI-crawler policy changes →

Work email only. No newsletter, no obligation.

Prefer to just buy it? See the sealed-archive monitoring offers, priced, with a dated sample.

The part nobody can go back and get

Every one of these files is free to fetch today. None of them is fetchable as it stood last week — sites overwrite robots.txt in place and publish no history. Across 55 sealed dates we have recorded 174,288 robots.txt body changes across 18,627 distinct domains. Most of those edits do not touch AI crawlers at all: 4,346 of them changed what at least one tracked AI crawler is allowed to do, across 2,793 domains and 31,194 individual crawler-verdict transitions. Between 2026-08-06 and 2026-08-07 (1 day), 3,206 of 61,057 compared domains changed their file and 106 changed an AI-crawler verdict.

See the full change history →

By crawler operator

An operator's tokens are governed independently: a site can block one and allow another.

By domain suffix

126 suffixes have at least 25 panel domains AND at least 25 of those serving a robots.txt we could parse, and get their own page. Suffix means the final DNS label, not the effective top-level domain: a .co.uk site is counted under .uk.

Declined, with counts (panel domains/parsed): 535 further suffixes appear in the panel and get no page, because a percentage over a handful of sites is noise — either too few sites carry the suffix, or too few of them answered at all. The second is a gap in our observation, not a finding about those sites. They are not hidden: .arpa (107/0), .click (68/22), .network (66/23), .win (61/24), .eg (52/24), .host (52/20), .ve (45/18), .sh (43/23), .asia (41/21), .services (40/6), .today (39/22), .digital (36/12), .lk (36/24), .art (35/20), .zone (35/12), .cyou (34/15), .google (34/24), .do (33/23), .name (33/16), .ms (32/15), .so (32/22), .work (32/12), .cam (31/22), .ink (31/15), .lol (31/15), .ac (30/19), .la (30/21), .li (30/21), .tz (30/16), .ba (29/24), .cr (29/20), .run (29/18), .team (29/11), .global (28/14), .systems (28/5), .website (28/11), .dz (27/6), .py (27/13), .kg (26/14), .mil (26/5), .sbs (26/11), .tn (26/11), .wiki (26/22), .int (25/20), .np (25/16), .plus (25/10), .cfd (24/9), .game (24/19), .icu (24/10), .pw (24/3), .lu (23/18), .tools (23/9), .gt (22/18), .page (22/9), .social (22/18), .bz (21/12), .xn--p1ai (21/10), .bid (20/12), .cx (20/9), .mz (20/12), .jo (19/8), .moe (19/13), .nu (19/10), .pub (19/9), .qpon (19/3), .science (19/8), .ug (19/12), .rocks (18/4), .email (17/7), .group (17/7), .guru (17/13), .hn (17/10), .ltd (17/9), .mk (17/13), .aero (16/8), .baby (16/7), .casino (16/8), .love (16/8), .ninja (16/6), .bo (15/10), .red (15/11), .ag (14/8), .bio (14/10), .cu (14/4), .fit (14/9), .goog (14/2), .onl (14/10), .technology (14/5), .aws (13/6), .best (13/9), .et (13/7), .mn (13/6), .pics (13/8), .rest (13/4), .sc (13/9), .stream (13/7), .sv (13/9), .wtf (13/10), .buzz (12/9), .ci (12/8), .cv (12/7), .eus (12/7), .gd (12/5), .hosting (12/8), .market (12/8), .ni (12/6), .re (12/5), .sex (12/12), .sx (12/7), .travel (12/8), .codes (11/5), .cool (11/5), .cy (11/8), .direct (11/4), .domains (11/4), .fan (11/6), .gh (11/6), .gs (11/6), .lat (11/7), .microsoft (11/1), .now (11/9), .pa (11/9), .party (11/9), .press (11/9), .qa (11/4), .studio (11/9), .trade (11/7), .uno (11/5), .watch (11/9), .zw (11/8), .al (10/5), .center (10/4), .company (10/4), .day (10/9), .inc (10/7), .lb (10/7), .men (10/6), .ps (10/5), .zm (10/4), .bar (9/6), .blue (9/4), .cm (9/6), .coop (9/2), .delivery (9/5), .events (9/3), .mg (9/6), .ml (9/5), .money (9/7), .mu (9/6), .ooo (9/4), .rw (9/7), .vc (9/6), .works (9/3), .ao (8/4), .cafe (8/6), .cash (8/5), .city (8/5), .date (8/6), .design (8/6), .fans (8/7), .gl (8/6), .gold (8/4), .help (8/7), .iq (8/7), .menu (8/8), .mom (8/3), .mt (8/4), .mv (8/3), .na (8/3), .om (8/5), .rodeo (8/6), .solutions (8/2), .tt (8/6), .amazon (7/2), .build (7/2), .care (7/6), .finance (7/3), .fo (7/3), .forum (7/3), .fyi (7/4), .ga (7/4), .golf (7/0), .hot (7/6), .kim (7/4), .kw (7/4), .mov (7/5), .ovh (7/2), .security (7/4), .support (7/2), .tj (7/4), .tl (7/4), .ad (6/3), .audio (6/3), .bh (6/3), .bj (6/6), .dog (6/3), .download (6/3), .engineering (6/3), .expert (6/3), .gy (6/4), .land (6/4), .mm (6/2), .museum (6/5), .pm (6/5), .rip (6/3), .sn (6/4), .software (6/4), .tg (6/1), .tm (6/4), .vg (6/2), .you (6/3), .bond (5/2), .box (5/4), .bw (5/3), .canon (5/3), .cards (5/2), .cd (5/3), .dating (5/2), .desi (5/5), .energy (5/2), .fox (5/0), .jobs (5/3), .london (5/2), .monster (5/2), .movie (5/4), .mw (5/5), .nf (5/2), .pet (5/2), .pizza (5/3), .rent (5/5), .sb (5/5), .scot (5/1), .tc (5/2), .tk (5/3), .wf (5/2), .academy (4/3), .apple (4/0), .as (4/3), .auction (4/2), .autos (4/2), .bike (4/4), .bingo (4/1), .bmw (4/0), .bnpparibas (4/0), .bot (4/3), .broker (4/1), .bt (4/3), .business (4/1), .casa (4/2), .community (4/3), .courses (4/2), .education (4/2), .exchange (4/4), .family (4/3), .farm (4/2), .film (4/3), .garden (4/3), .gle (4/1), .green (4/2), .hair (4/3), .health (4/3), .ht (4/2), .lc (4/2), .lifestyle (4/1), .loan (4/2), .mo (4/3), .mp (4/2), .nexus (4/2), .parts (4/2), .pink (4/2), .place (4/2), .quest (4/1), .report (4/4), .sale (4/2), .skin (4/4), .style (4/3), .tf (4/3), .tokyo (4/2), .town (4/4), .va (4/2), .vet (4/4), .vu (4/2), .wales (4/1), .webcam (4/4), .ye (4/0), .africa (3/2), .agency (3/0), .army (3/0), .beauty (3/2), .beer (3/0), .bi (3/3), .black (3/0), .camera (3/0), .capital (3/1), .cf (3/1), .coffee (3/1), .dad (3/3), .express (3/2), .faith (3/1), .fj (3/2), .foo (3/2), .foundation (3/2), .frl (3/1), .gal (3/2), .gift (3/2), .godaddy (3/0), .horse (3/1), .house (3/1), .ing (3/1), .international (3/1), .istanbul (3/2), .je (3/3), .ki (3/3), .law (3/2), .llc (3/1), .ls (3/2), .mba (3/2), .nrw (3/2), .partners (3/2), .photos (3/2), .pictures (3/1), .promo (3/1), .review (3/1), .sarl (3/3), .school (3/2), .show (3/2), .sky (3/0), .tel (3/1), .tips (3/2), .vin (3/0), .vision (3/3), .yandex (3/2), .youtube (3/2), .abbott (2/2), .actor (2/0), .auto (2/1), .azure (2/0), .bank (2/1), .basketball (2/2), .berlin (2/1), .bf (2/2), .bn (2/1), .boats (2/1), .bradesco (2/1), .bs (2/1), .cab (2/1), .christmas (2/1), .church (2/1), .ck (2/1), .clinic (2/1), .credit (2/0), .cw (2/1), .cymru (2/2), .deals (2/2), .diet (2/2), .diy (2/2), .earth (2/2), .forsale (2/0), .free (2/2), .fund (2/1), .gay (2/0), .globo (2/0), .gm (2/2), .gratis (2/1), .gripe (2/0), .homes (2/1), .industries (2/1), .institute (2/2), .jm (2/1), .kh (2/2), .kpmg (2/2), .krd (2/1), .leclerc (2/1), .lr (2/0), .management (2/0), .markets (2/2), .motorcycles (2/2), .navy (2/0), .ngo (2/0), .nhk (2/0), .nr (2/1), .nyc (2/1), .photo (2/2), .pn (2/2), .pr (2/2), .realtor (2/2), .restaurant (2/1), .rio (2/1), .sap (2/1), .sbi (2/0), .select (2/0), .shopping (2/0), .sm (2/1), .sncf (2/1), .sport (2/2), .spot (2/1), .sr (2/1), .study (2/1), .supply (2/1), .sy (2/1), .sz (2/2), .taxi (2/0), .toyota (2/0), .toys (2/1), .williamhill (2/0), .xin (2/0), .zip (2/1), .abb (1/1), .abudhabi (1/1), .adult (1/1), .af (1/1), .archi (1/0), .audi (1/0), .aw (1/0), .ax (1/1), .barcelona (1/1), .barclays (1/1), .bargains (1/1), .bb (1/0), .bbva (1/1), .bible (1/1), .boo (1/1), .boston (1/1), .brussels (1/1), .bzh (1/1), .car (1/0), .careers (1/1), .cars (1/0), .catering (1/0), .ceo (1/0), .cern (1/1), .cg (1/1), .channel (1/0), .chrome (1/1), .claims (1/0), .clothing (1/1), .computer (1/0), .cooking (1/0), .deal (1/0), .dhl (1/0), .dj (1/1), .dm (1/1), .eco (1/1), .edeka (1/1), .engineer (1/0), .fast (1/0), .feedback (1/1), .florist (1/0), .food (1/0), .football (1/1), .forex (1/1), .fujitsu (1/0), .gallery (1/0), .gent (1/1), .gi (1/1), .gives (1/0), .giving (1/0), .glass (1/1), .gn (1/0), .gp (1/1), .graphics (1/0), .guide (1/1), .gw (1/0), .hbo (1/0), .healthcare (1/1), .holiday (1/0), .honda (1/0), .how (1/1), .immo (1/0), .insure (1/1), .ist (1/1), .kitchen (1/1), .km (1/1), .koeln (1/1), .legal (1/1), .lgbt (1/0), .limo (1/1), .luxury (1/1), .madrid (1/0), .makeup (1/1), .marketing (1/0), .mc (1/0), .med (1/0), .moda (1/1), .moi (1/1), .nc (1/1), .ne (1/1), .nec (1/0), .new (1/1), .ntt (1/1), .observer (1/1), .office (1/0), .paris (1/1), .pf (1/0), .pg (1/1), .philips (1/1), .pictet (1/0), .pioneer (1/1), .racing (1/1), .ren (1/0), .republican (1/0), .reviews (1/0), .rugby (1/1), .saxo (1/1), .schwarz (1/0), .sd (1/0), .sexy (1/1), .sharp (1/1), .shiksha (1/1), .sl (1/1), .solar (1/1), .sony (1/0), .statefarm (1/0), .surf (1/0), .taipei (1/0), .talk (1/1), .tatar (1/0), .td (1/1), .theater (1/0), .toshiba (1/0), .training (1/1), .university (1/0), .uol (1/1), .vegas (1/1), .vi (1/1), .vodka (1/1), .wang (1/0), .wedding (1/0), .xn--80asehdb (1/1), .xn--c1avg (1/0), .xn--d1acj3b (1/0), .xn--p1acf (1/1), .xn--q9jyb4c (1/0), .yachts (1/1), .yoga (1/1).

Read this before quoting a trend. We hold 55 sealed daily snapshots between 2026-06-09 and 2026-08-07, out of 60 calendar days. 5 calendar day(s) inside that span have no collection run at all: 2026-06-11, 2026-06-12, 2026-06-19, 2026-07-04, 2026-07-15. Consecutive rows in every table below are consecutive SEALED dates, not consecutive days — where a gap falls, that row's diff covers more than 24 hours, and the span is printed with it. We do not interpolate across a day we did not collect.
Why you will not find a robots.txt file here. These pages render mechanical verdicts and aggregate diffs, never a fetched body. Site operators leave contact addresses in robots.txt comments — on the 2026-08-06 snapshot 2,200 of the distinct bodies we fetched contain an email-shaped string — and republishing bodies would republish those. Nothing on this surface identifies an individual site, either: every figure is a count.

Panel: 100,000 domains, frozen research list (tranco_L8QG4_top100000.txt:8778484edc1487aa, source: Tranco). We publish our own dated observations over that panel; no Tranco rank is republished. Methodology census-v1. Full methodology, including the transport-error taxonomy and what a 200 does not mean.

Compiled from each site's own published /robots.txt, /ai.txt, /llms.txt and /.well-known/tdmrep.json, fetched daily over a frozen 100,000-domain research panel (source list: Tranco; no Tranco rank is republished here). Verdicts are mechanical readings of what those files say under RFC 9309 grouping with exact product-token matching — they are not legal advice, not a statement about any site's intentions, and not a claim about what any crawler actually did. A robots.txt is a request, not an access control. Figures are provided for informational purposes only and carry no warranty of accuracy or completeness.

AI crawler policy home · What changed · How this is collected · US Tech Automations.