19.3% of Top Sites Block AI Crawlers: USTA Index
The US Tech Automations Closing Web Index found that 10573 sites—19.3% of the parseable set—block at least one tracked AI crawler. This is point-in-time AI-crawler access policy for the Tranco top 100,000 domains, parsed from public robots.txt collected on August 1, 2026; 54,730 returned a parseable robots.txt, and every percentage uses that parseable set rather than all domains attempted.
The edition is cross-sectional. It does not say whether blocking rose or fell, and it does not represent the whole web.
Key Findings
10573 sites—19.3%—block at least one tracked AI crawler. The sealed Closing Web snapshot counts only exact crawler groups containing Disallow: /.
GPTBot is blocked by 8507 sites, or 15.5%. That is the largest per-bot count among the crawler rows reported in this edition.
OpenAI is blocked by 8601 sites, or 15.7%. Operator totals combine the named user-agents assigned to an operator and count each site once.
2175 sites carry a wildcard block, or 4%. A wildcard block is a separate signal from a crawler-specific rule.
18510 /llms.txt files were observed across the full crawl. Within the parseable-robots subset, the sealed adoption rate is 24.2%; those fields have different universes and must not be treated as one fraction.
The headline is selective access policy, not universal closure. The snapshot tracks 21 bot identifiers tied to 12 operators, and the reported bot rates range from GPTBot at 15.5% to PerplexityBot at 4.4% among the listed rows.
The Index at a Glance
| Sealed measure | Value | Denominator or scope |
|---|---|---|
| Domains attempted | 100000 | Tranco top 100,000 |
| Parseable robots.txt files | 54730 | Rate denominator |
| Bots tracked | 21 | Exact user-agent identifiers |
| Operators tracked | 12 | Grouped bot ownership |
| Block at least one tracked crawler | 10573 | 19.3% of parseable robots.txt files |
| Wildcard block | 2175 | 4% of parseable robots.txt files |
| /llms.txt files observed | 18510 | Full 100000-domain crawl |
| /llms.txt adoption rate | 24.2% | Parseable-robots subset |
| Snapshot date | August 1, 2026 | Point in time |
The denominator column is part of the result. The block rates describe only domains whose robots.txt could be parsed. A domain without a reachable or parseable file is neither added to a block count nor silently categorized as allowing a crawler.
The /llms.txt fields require another distinction. The count records observations across the full crawl, while the rate is calculated inside the parseable-robots subset. Reporting both is useful only when their universes remain visible.
Which AI Crawlers Get Blocked Most
| Crawler | Sites blocking | Share of 54730 |
|---|---|---|
| GPTBot | 8507 | 15.5% |
| CCBot | 8281 | 15.1% |
| Bytespider | 7802 | 14.3% |
| ClaudeBot | 7541 | 13.8% |
| Google-Extended | 6976 | 12.7% |
| Amazonbot | 6843 | 12.5% |
| Meta-ExternalAgent | 6643 | 12.1% |
| Applebot-Extended | 6533 | 11.9% |
| PerplexityBot | 2387 | 4.4% |
GPTBot and CCBot sit at the top of the reported per-bot table. Bytespider and ClaudeBot follow, while Google-Extended, Amazonbot, Meta-ExternalAgent, and Applebot-Extended form the next visible group. PerplexityBot is lower than the other listed bot rows.
This ordering describes exact text found in public robots.txt files. It does not establish why a webmaster wrote a rule, whether the crawler obeyed it, how often the bot visited, or whether another control also limits access.
A crawler name may also differ from its operator total. The per-bot table asks whether one exact user-agent token has a full disallow. The operator table asks whether any token assigned to that operator has a full disallow, then counts the site once.
The Operator-Level Picture
| Operator | Sites blocking | Share of 54730 |
|---|---|---|
| OpenAI | 8601 | 15.7% |
| Common Crawl | 8281 | 15.1% |
| Anthropic | 7888 | 14.4% |
| ByteDance | 7802 | 14.3% |
| Meta | 7196 | 13.1% |
| 6976 | 12.7% | |
| Amazon | 6843 | 12.5% |
| Apple | 6533 | 11.9% |
| Perplexity | 2400 | 4.4% |
| Cohere | 2315 | 4.2% |
| Diffbot | 2152 | 3.9% |
| Mistral | 609 | 1.1% |
OpenAI has the largest operator-level count in this edition, followed by Common Crawl, Anthropic, and ByteDance. Mistral has the smallest operator row. Those positions are a description of this sealed table, not a general measure of trust, traffic, market share, or compliance.
OpenAI's operator figure is above the GPTBot row because its group includes GPTBot, OAI-SearchBot, and ChatGPT-User. Anthropic's group includes ClaudeBot and other named Anthropic user-agents. Meta and Perplexity also have more than one token in their operator groups.
Google, Common Crawl, ByteDance, Amazon, Apple, Cohere, Mistral, and Diffbot have operator totals that align with the exact tokens assigned in this snapshot. Alignment does not imply a permanent relationship; it reflects the grouping sealed for this edition.
How to Read robots.txt and /llms.txt Together
A full crawler block means an exact user-agent group contains
Disallow: /in the domain's own robots.txt. Partial-path rules do not qualify for this index.
An
/llms.txtfile is an observed policy or guidance artifact, not an enforcement mechanism. Its presence does not prove that a named crawler is allowed, blocked, or compliant.
The two files answer different questions. robots.txt expresses crawl directives by user-agent and path. /llms.txt can publish machine-readable guidance for AI systems. A site may use either, both, or neither; this snapshot does not infer intent from the combination.
The July Closing Web Index is a separate sealed edition. It can provide historical context, but this page makes no change claim because it reports only the August 1 cross-section.
The table also should not be converted into an “allow” list. A missing exact full block can coexist with partial restrictions, a wildcard rule with overrides, authentication, paywalls, application controls, or policies outside robots.txt. The index measures its stated syntax and nothing broader.
What This Snapshot Can Support
The snapshot can support a direct answer about the named ranking on the named date: whether a parseable robots.txt contains an exact full block for a tracked crawler token. It can also support side-by-side reading of bot and operator rows because those rates share the same parseable denominator.
The snapshot can support an audit queue. A team may use the sealed result to decide which domains deserve a current policy check before collection, retrieval, or AI-visibility analysis. The live file should still be fetched again at the moment of action because this edition is intentionally point-in-time.
The snapshot cannot support a conclusion about webmaster motivation. A rule may reflect licensing, load, privacy, security, search policy, a template, or another concern. The file contains the directive, not the meeting or policy process that produced it.
It also cannot prove crawler compliance. The relevant evidence for behavior would come from server, CDN, or application logs tied to the crawler and request. This index observes public instructions and keeps its claim at that layer.
Finally, the index cannot rank the entire internet. The Tranco list is a defined high-traffic ranking, which makes the census repeatable, but it is not every hosted domain and is not presented as a random sample from a larger population.
Methodology
Source attribution: Closing Web crawl — Tranco top 100,000 domains, using public robots.txt and /llms.txt, collected and sealed point-in-time.
Every numeral is a verbatim count computed from robots.txt files the closing-web clock actually fetched on the snapshot date. “Blocks” means an exact user-agent group containing Disallow: /. The universe is the Tranco top-100K ranking; percentages are over sites returning a parseable robots.txt, not all domains attempted. Domains without a reachable or parseable robots.txt are excluded from rate denominators.
In this report, nothing is estimated, modeled, or extrapolated. The snapshot is a census of the named Tranco ranking and the public files returned to the clock, not a random sample, a projection about the whole web, or evidence that any crawler obeyed a directive.
Fetch each ranked domain's public robots.txt and
/llms.txtat the snapshot clock.Exclude robots.txt responses that cannot be reached and parsed from every rate denominator.
Parse exact user-agent groups and count only groups containing
Disallow: /for the named crawler.Group bot tokens by the sealed operator mapping, deduplicate each domain inside an operator total, and seal the snapshot with its content hash.
The method intentionally refuses semantic guesses. It does not count a partial restriction as a full block, infer policy from a missing file, or translate a broad statement elsewhere on the site into a robots.txt decision.
The July crawler-movement report uses a separately sealed derived method for its window. That is the appropriate page for movement analysis; this edition remains cross-sectional.
Frequently Asked Questions
Q: What counts as blocking an AI crawler?
A: The domain's own robots.txt must contain a user-agent group that exactly names the tracked crawler and includes Disallow: /. A partial path restriction is not counted as a full block.
Q: Why are block percentages based on 54730 sites?
A: Those sites returned a parseable robots.txt. Domains without a reachable and parseable file are excluded rather than treated as either blocking or allowing a crawler.
Q: Does a robots.txt disallow force a crawler to stop?
A: The index observes published directives, not crawler behavior. Whether a bot requests a page or follows the file requires different evidence, such as server-side logs and the operator's documented behavior.
Q: What does the 24.2% /llms.txt rate mean?
A: It is the sealed adoption rate within the parseable-robots subset. The separate 18510 count records /llms.txt observations across the full 100000-domain crawl, so the count and rate must keep their stated universes.
Q: Does this cover the whole web?
A: No. It covers the Tranco top 100,000 ranking on August 1, 2026. It is not a census of every domain and should not be projected beyond the named ranking.
Q: Why can an operator total differ from a bot total?
A: An operator may have several mapped user-agent tokens. The operator total counts a domain once when any mapped token has the exact full block, while a bot row tests only its named token.
Put AI-Access Data to Work
An AI-visibility or GEO lead can re-crawl a client corpus and route a changed access rule to the content owner before a retrieval plan assumes the source remains available. A publisher RevOps lead can monitor owned domains and send policy changes through editorial, legal, and revenue review.
A competitive-intelligence analyst or retrieval team can maintain a permitted-source register, recheck the highest-value domains on a defined cadence, and pause downstream collection when an exact full block appears. The business-site AI-access benchmark adds a narrower business-domain lens to that operating decision.
US Tech Automations can automate the recurring work: fetch the approved domain set, parse exact directives, compare the newly sealed state with the prior authorized state, alert the named owner, and preserve the evidence that changed the retrieval rule.
Use agentic workflow monitoring to turn crawler-policy observations into owned review and escalation instead of an occasional spreadsheet check.
Source: US Tech Automations Research — computed from the sealed Closing Web snapshot, August 1, 2026.
Get this data as a daily feed
The numbers in this report come from a permit feed we monitor daily. Leave your email and we will follow up about a daily feed for your ZIPs and categories.
Prefer to talk first? Contact us.
Cite this report
US Tech Automations Research, 2026-08 edition. “19.3% of Top Sites Block AI Crawlers: USTA Index.” https://ustechautomations.com/resources/blog/closing-web-index-august-2026
Sealed snapshot sha256: 69544d238343330cf3a3403f61913938ad63a90779ba99e4bcc6c4bd4b82938e
Machine-readable data: CSV · JSON · All research & methodology
About the Author

Helping businesses leverage automation for operational efficiency.