Research & Data

AI-Crawler Blocking Climbed to 19.4% on Top Sites

Aug 8, 2026

AI-crawler blocking reached 19.4% of the measured sites on August 7, 2026. The same measure was 18.5% on June 9, 2026, an exact +0.9 percentage-point movement across the two sealed endpoints. This is a useful access-policy signal, but it is not a verdict on AI visibility, crawler behavior, or the web as a whole.

Change in AI-crawler access policy across the Tranco top 100,000 between two sealed closing-web editions: June 9, 2026 and August 7, 2026. Every endpoint is a verbatim count from its own sealed edition; every delta is the arithmetic difference of those two sealed editions, in percentage points.

The relevant denominator is not every domain in the ranking. It is the sites where the collection produced a parseable robots.txt file: 54904 at the first endpoint and 54927 at the later endpoint. That distinction makes the result a change in share, not a count of newly blocked sites. The six editions in the window provide repeat observations, but the window is still only 59 days long.

19.4% is the August endpoint for any tracked AI-crawler block in the measured robots.txt population.

The +0.9 movement is in percentage points, not a claim about raw domains or all websites.

The short answer for an AI-visibility team

A robots.txt block is a public instruction that requests certain automated clients not fetch specified paths. It is a policy signal from a website operator. It does not prove that every crawler obeyed, that every page is unavailable to every system, or that a site will disappear from any answer engine. The direct answer is therefore narrow: the share of parseable robots.txt files carrying an any-AI-crawler block was higher at the later sealed endpoint.

MeasureJune 9, 2026August 7, 2026Sealed movement
Parseable robots.txt sites5490454927not a common raw denominator
Any tracked AI-crawler block18.5%19.4%+0.9
llms.txt observed22.6%24.5%+1.9
Tracked crawler policies21210

The sample is the Tranco top 100,000, not a census of the internet. It also is not an audit of page rendering, authentication walls, JavaScript behavior, or a crawler request log. It measures the public policy artifacts collected from that sample. Readers deciding whether to change a site policy should use this as context and inspect their own directives, server behavior, and product requirements.

The earlier Closing Web June edition and the July edition supply useful point-in-time context. This report is different because it fixes the comparison to two named endpoints and preserves their separate denominators.

Which crawler policies moved

The broad measure rose, but it would be wrong to say every tracked crawler rose. Of 21 tracked crawler policies, 19 were blocked by a larger share of the measured sites, 2 by a smaller share, and 0 were unchanged. The table names a few representative policy measures rather than treating a family of crawlers as interchangeable.

Public policy measureJune 9, 2026August 7, 2026Change
GPTBot block share14.9%15.6%+0.7
ClaudeBot block share12.9%13.8%+0.9
Google-Extended block share12%12.9%+0.9
CCBot block share14.4%15.2%+0.8
PerplexityBot block share4.4%4.3%-0.1

Those labels identify instructions observed in robots.txt. They do not establish a shared owner, a shared purpose, or equivalent crawler behavior. A publisher might set one directive and not another for reasons that have nothing to do with answer visibility. Similarly, a smaller share for one policy does not negate the broader movement.

GPTBot policy blocks moved from 14.9% to 15.6%.

ClaudeBot policy blocks moved from 12.9% to 13.8%.

The any-AI-policy share moved from 18.5% to 19.4%.

For definitions of crawler behavior, consult the documented policies for OpenAI bots, Anthropic crawling, and Google common crawlers. Those documents explain particular products; they do not turn this sample into a compliance test.

Operator patterns are easier to read than a single headline

Grouping observed policy labels by operator can help an SEO director see whether a single broad headline hides divergent measures. The values below remain policy shares, not traffic shares and not crawl counts.

Policy groupJune 9, 2026August 7, 2026Change
OpenAI15%15.8%+0.8
Anthropic13.6%14.5%+0.9
Google12%12.9%+0.9
Meta12.1%13.3%+1.2
ByteDance13.2%14.4%+1.2

This view answers a practical monitoring question: are public directives changing in one measured population? It cannot answer whether a particular model cited a particular page, whether a policy was honored, or why a site owner changed it. Causation is outside the snapshot.

The companion Closing Web Index provides the original index framing. A recurring trend page should not replace a site-level audit: directives can be scoped by path, can be overridden by product-specific controls, and may change between crawls.

Why denominator discipline matters

The first collection found 54904 parseable robots.txt files; the later collection found 54927. Reporting a percentage-point change keeps those endpoints comparable without pretending they have the same raw population. Converting the percentages into a new estimated site count would add a calculation that the sealed record does not publish, so this article does not do that.

The index also tracks llms.txt observations, which rose from 22.6% to 24.5%. An llms.txt file is neither a universal crawler permission nor a replacement for robots.txt. Its presence does not prove indexing, retrieval, or a model response. Treat it as a separately observed public artifact.

This is a derived trend seal. Endpoint percentages are copied verbatim from the two sealed base editions named in the source line; deltas are the exact subtraction of those two figures, in percentage points, nothing modeled or extrapolated. Each edition's denominator is the sites that returned a parseable robots.txt that day, so a delta is a change in share, not a change in raw site count. Scope is public robots.txt for the Tranco ranking only, not the whole web. Nothing is estimated, modeled, or extrapolated.

Put AI-visibility monitoring to Work

For a content lead, retrieval team, or SEO manager, the operational question is usually not whether a benchmark moved. It is whether the organization has an intentional, reviewable policy. A sensible workflow records the current directive, the intended policy owner, the pages or paths affected, and a change-review date. It also separates discovery, training, and product-specific access questions instead of applying a blanket rule by habit.

US Tech Automations can help an SEO or content operations lead automate that review trail: collect approved public policy files, detect a change, route it to the named owner, and retain the decision with the affected site section. The workflow is about reliable review and escalation, not a promise of rankings, inclusion, or crawler compliance. See how that pattern fits into agentic workflows when a team needs an auditable handoff rather than another dashboard.

Avoid using benchmark movement as a trigger for an unreviewed sitewide block or allow rule. A policy can affect legitimate tools as well as unwanted ones. Start with the organization’s purpose, confirm the actual directive syntax, and test changes in the environment where they will operate.

Frequently Asked Questions

Q: Does 19.4% mean that 19.4% of the web blocks AI?

A: No. It is the share of parseable robots.txt files observed for the Tranco top 100,000 sample at the August 7, 2026 endpoint. The ranking sample and the parseable-file denominator are both narrower than the web.

Q: Does the +0.9 change mean more sites were added to a block list?

A: No. It is a percentage-point difference between two shares whose parseable-robots.txt denominators differ. The seal does not publish a raw additions list or a reason for each policy change.

Q: Does a robots.txt policy prove a crawler did not fetch a page?

A: No. The snapshot records public access-policy text, not crawler logs, enforcement outcomes, or page-level fetching behavior.

Q: Is llms.txt a substitute for robots.txt?

A: No. The observed llms.txt share is a separate artifact. Its presence does not prove permission, indexing, retrieval, or a response outcome for any AI system.

Q: Should a company change policy because the benchmark rose?

A: Not automatically. The useful next step is a policy review by the content, SEO, legal, and platform owners who understand the site’s goals and the consequences of each directive.

Source: US Tech Automations Research, Closing Web Index trend seal, June 9, 2026 – August 7, 2026. Public robots.txt and llms.txt observations from a fixed Tranco sample; aggregate access-policy facts only.

How to read an access-policy trend responsibly

The strongest use of this report is comparative. A site owner can ask whether public directives in a fixed, high-visibility sample are moving in one direction over a short, named interval. That is materially different from asking whether a given AI product has indexed a page, whether a company should allow a product, or whether an answer engine will cite a domain. Each of those questions requires a different source of evidence.

Start with the artifact that the report actually observes: a robots.txt response that can be parsed at crawl time. A parseable response is useful because it allows the same policy fields to be evaluated across the sample. It is also a limitation. Sites without a parseable response are not silently converted into a policy position, and a response can change after the capture. The denominator rule keeps that uncertainty visible.

An AI-visibility review should also separate the intended audience for a directive. A marketing team may care about discovery in public answer systems. A platform team may care about resource usage, security, or preserving a service boundary. A legal or editorial team may care about rights and attribution. Those are legitimate concerns, but they should not be collapsed into a universal “block” or “allow” label without an owner, a reason, and a method for review.

The operator breakdown is helpful precisely because it resists that collapse. A website can expose different policies to different named clients. A measured difference between public directives does not tell us whether the clients share a crawl cadence, a retention policy, or a downstream use. It only tells us that the site-level policy strings differed in the observed population. That is enough for a monitoring dashboard; it is not enough for an enforcement conclusion.

For recurring reporting, preserve the two endpoints, the denominator at each endpoint, the sampled universe, and the policy definition beside the headline. If any one of those changes, a time series can become visually persuasive but analytically misleading. This publication keeps those elements in the body and the downloadable asset so a reader does not have to reverse-engineer them from a chart.

The same restraint applies to business decisions. A benchmark is best used to prompt a documented review: what does our current directive say, which public agents does it address, which business purpose does it serve, and who can approve a change? The report cannot answer those organization-specific questions. It can make them harder to ignore.

An additional safeguard is to document the paths and environments a directive is intended to cover. A public benchmark can provide context, but it cannot decide an organization-specific policy. For clarity, nothing is estimated, modeled, or extrapolated beyond the sealed policy observations in this report.

Get this data as a daily feed

The numbers in this report come from a permit feed we monitor daily. Leave your email and we will follow up about a daily feed for your ZIPs and categories.

Prefer to talk first? Contact us.

Cite this report

US Tech Automations Research, 2026-08 edition. “AI-Crawler Blocking Climbed to 19.4% on Top Sites.” https://ustechautomations.com/resources/blog/ai-crawler-blocking-trend-august-2026

Sealed snapshot sha256: e809e09a9e0ecf5942c0a80a7b04f0a2d209ad8026b733071f3b48a77d988f44

Machine-readable data: CSV · JSON · All research & methodology