AI Crawler Blocking by Bot: Weekly Closing Web Index
The weekly Closing Web Index asks a deliberately narrow question: for a named crawler requesting the root path /, what did each stored robots.txt policy say on August 11, 2026? The source observation cutoff is 2026-08-11. The answer is split into three buckets so missing evidence cannot masquerade as permission.
For GPTBot, 10,214 frozen-panel domains had an applicable rule that disallowed /. Another 59,028 were not disallowed by the tested rule, while 30,758 remained UNKNOWN. The classified-policy denominator was therefore 69,242, not the full 100,000-domain panel.
GPTBot's classified-policy disallow rate was 14.8%.
Transport-error missingness is systematic; classified-policy rates describe the answered subpopulation, while panel and UNKNOWN shares use the full frozen panel.
UNKNOWN covered 30.8% of the full panel.
Those two statements belong together. The first describes the policies the read model could classify. The second describes the evidence it could not classify. Neither is a claim about crawler obedience, indexing, model training, answer-engine citations, or legal authorization.
This is an in-place weekly upgrade of the existing August research URL. It replaces the page's older shortcut parser and denominator with a method-frozen root-path test that can be rebuilt from the append-only seal. US Tech Automations Research publishes the aggregate result, its method, and its hashes; it does not publish the stored policy bodies or attach verdicts to named domains.
The latest named-crawler scoreboard
The table below covers GPTBot, ClaudeBot, CCBot, Google-Extended, and PerplexityBot. The classified denominator keeps the evidence gap visible beside every positive finding.
| Crawler | DISALLOWED | Not disallowed by tested rule | UNKNOWN | Classified denominator |
|---|---|---|---|---|
| GPTBot | 10,214 | 59,028 | 30,758 | 69,242 |
| ClaudeBot | 9,342 | 59,900 | 30,758 | 69,242 |
| CCBot | 10,127 | 59,115 | 30,758 | 69,242 |
| Google-Extended | 8,729 | 60,513 | 30,758 | 69,242 |
| PerplexityBot | 4,140 | 65,102 | 30,758 | 69,242 |
The classified-policy rates were GPTBot 14.8%, ClaudeBot 13.5%, CCBot 14.6%, Google-Extended 12.6%, and PerplexityBot 6%; their full-panel shares were 10.2%, 9.3%, 10.1%, 8.7%, and 4.1%. Transport-error missingness is systematic; classified-policy rates describe the answered subpopulation, while panel and UNKNOWN shares use the full frozen panel.
The panel share is smaller than the known rate because it keeps UNKNOWN domains in the denominator. A reader can use the known rate to compare classified policies and the panel share to understand how much of the complete frozen universe produced a positive disallow finding.
“Not disallowed by tested rule” is intentionally not shortened to “allowed.” It means only that the read model did not find an effective Disallow for the tested token and path. A different path, a different crawler token, a control outside robots.txt, or a later policy could yield another result.
The June Closing Web Index remains useful as the original cross-sectional framing. This weekly page now owns the recurring root-path measure, so readers do not have to reconcile another near-duplicate URL.
What each verdict means
The read model applies one frozen interpretation to both comparison dates. It does not infer intent from a company name, a comment, or a policy's surrounding prose.
| Verdict | Mechanical meaning at / | What it does not mean |
|---|---|---|
| DISALLOWED | The applicable group’s winning rule is Disallow | The crawler obeyed, the content vanished, or access control exists |
| NOT_DISALLOWED_BY_TESTED_RULE | No applicable winning Disallow was observed for this token and path | The whole site is open, every path is crawlable, or access is authorized |
| UNKNOWN | The stored observation could not support a trustworthy rule verdict | The site allowed or blocked the crawler |
Group selection follows the Robots Exclusion Protocol model. Product-token matches are exact and case-insensitive. When the named token has one or more exact groups, those groups are merged and the wildcard group does not replace them. The wildcard group applies only when no exact group exists.
Rule selection is path-specific. The longest matching rule wins, and Allow wins an equal-length tie. An empty Disallow does not block. Other records such as Sitemap do not silently terminate a user-agent group. These details matter because merely searching a file for Disallow: / can count the wrong group or ignore a more specific Allow.
The build also runs a discriminating negative control. A fixed fixture gives GPTBot a root Disallow and a fabricated token a root Allow. The two verdicts must differ before the live seal is read. That check makes a parser that ignores the requested token fail instead of passing on an empty or constant result.
What changed during the week
The comparison uses August 4, 2026 and August 11, 2026, exactly one week apart on the same frozen panel and collection method. The known-policy denominator changed from 69,583 to 69,242, so the table reports rates and their percentage-point movement rather than pretending both dates classify the same raw population.
The August 4 to August 11 classified-policy rates were GPTBot 14.7% → 14.8%, ClaudeBot 13.5% → 13.5%, CCBot 14.6% → 14.6%, Google-Extended 12.6% → 12.6%, and PerplexityBot 6% → 6%. Transport-error missingness is systematic; classified-policy rates describe the answered subpopulation, while panel and UNKNOWN shares use the full frozen panel.
The safe reading is that these aggregate root-policy rates were mostly stable across the week. GPTBot's rounded known rate moved by +0.1 percentage point. That does not establish that a particular domain added a rule, because the classified population also changed. The public artifact intentionally contains no domain-level transition list.
The July policy trend used the earlier edition method. It provides historical context, but its headline should not be spliced into this series as if the definitions were identical. This page begins the stricter weekly contract: named token, tested path, full partition, disclosed exclusions, and reproducible hash binding.
Why UNKNOWN is larger than transport failures
Transport errors were the largest unknown source, but they were not the only one. A response can arrive and still fail to support a trustworthy root-policy verdict. The read model therefore records each exclusion reason before calculating any crawler rate.
This is not random-looking attrition that can be hidden in a footnote. The preregistered BIASED verdict requires more than 5% of domains to fail on at least 90% of dates. Across the August 11 publication window, 23,102 of 100,000 domains (23.1%) had a transport error on at least 53 of 58 sealed dates, so the result crosses that bound. A domain-first query and an independent date-first query returned the same 23,102-domain set with zero disagreements. Transport-error missingness is systematic; classified-policy rates describe the answered subpopulation, while panel and UNKNOWN shares use the full frozen panel.
| Frozen Tranco rank band | Domains in band | Persistent-error domains |
|---|---|---|
| 1–1,000 | 1,000 | 309 |
| 1,001–10,000 | 9,000 | 2,132 |
| 10,001–25,000 | 15,000 | 3,846 |
| 25,001–50,000 | 25,000 | 6,551 |
| 50,001–100,000 | 50,000 | 10,264 |
The persistent shares by rank band were 30.9%, 23.69%, 25.64%, 26.2%, and 20.53%, respectively. Transport-error missingness is systematic; classified-policy rates describe the answered subpopulation, while panel and UNKNOWN shares use the full frozen panel.
| Latest UNKNOWN reason | Frozen-panel count |
|---|---|
| Transport error | 24,376 |
| Server error | 464 |
| Redirect or other status | 139 |
| Missing stored body | 701 |
| Truncated body | 400 |
| HTML-shaped soft response | 4,443 |
| Invalid UTF-8 | 235 |
| Total UNKNOWN | 30,758 |
A transport failure is not converted to “open,” “no policy,” or zero. A server error is also not interpreted as a publisher's durable preference. A truncated policy may omit the decisive rule, and an HTML application shell served from a robots URL is not treated as a clean policy file. Invalid text is kept outside the classified denominator rather than decoded with silent replacement characters.
Client-error responses receive the narrow NOT_DISALLOWED_BY_TESTED_RULE label because no tested rule was available in that dated robots response. Even there, the wording stops short of permission. Robots instructions are not access authorization, and a website may enforce controls elsewhere.
The August cross-sectional index is still available for its original monthly view. The weekly method on this page is the better source when the question is specifically about a named crawler at / and the size of the unknown population matters.
Coverage and integrity
The source contains 58 complete dates from June 9, 2026 through August 11, 2026. There are 6 calendar gaps in that span: June 11, June 12, June 19, July 4, July 15, and August 8. The page does not interpolate those dates or draw a smooth line through them.
A date is complete only when policy_snapshots contains one sealed row for every frozen-panel domain and each of the four collected resources. The count comes from the rows themselves, not from a run log. Both weekly endpoints contain 100,000 stored robots cells and belong to the same frozen Tranco panel and collection methodology. A stored cell is not automatically an observation.
Stored rows are not observations
Through August 11, the dated store contains 23,200,000 sealed grid cells. That is not an observation count. The fixed partition finds 17,557,013 HTTP answers and 5,642,987 transport failures; together they reconcile to every stored cell with a difference of 0. Any non-null HTTP status is an answer, including a 404. A row with no HTTP status and a non-null fetch_error observed no HTTP answer and stays outside the observation count.
| Counted statement | Exact row predicate (:publication_cutoff is August 11, 2026) |
|---|---|
| 23,200,000 stored grid cells are not observations | policy_snapshots.snapshot_date <= :publication_cutoff |
| 5,642,987 transport failures are not observations | policy_snapshots.snapshot_date <= :publication_cutoff AND status_code IS NULL AND fetch_error IS NOT NULL |
| 17,557,013 answered observations | policy_snapshots.snapshot_date <= :publication_cutoff AND status_code IS NOT NULL AND fetch_error IS NULL |
| 5,934,853 retained HTTP-200 bodies | policy_snapshots.snapshot_date <= :publication_cutoff AND status_code = 200 AND content_sha256 IS NOT NULL AND EXISTS (SELECT 1 FROM blobs WHERE blobs.content_sha256 = policy_snapshots.content_sha256) |
| 1,020,697 distinct retained blobs | blobs.first_seen_date <= :publication_cutoff |
The predicates are part of the downloadable JSON beside the counts. A nonzero partition difference, an ambiguous answer/error row, or an HTTP-200 hash without a retained blob is a finding; the query is not adjusted to preserve an earlier total.
Before an aggregate is emitted, the builder re-derives every selected robots row hash and every referenced body hash. It also recomputes the collection batch hash from the full sealed date. The public JSON is bound to those source hashes, the frozen universe file, the clock methodology, and the read-model code by artifact sha256 5ba00c28b989c6da9883d58a11770fbad7a5a435edb5f8c0b4ebd8d2729835e2.
The JSON and CSV twins contain aggregates and hashes only. They serialize no robots.txt body, response header, domain identity, or per-domain verdict. That boundary preserves the citable measurement without turning a public download into a list of claims about individual organizations.
For clarity, nothing is estimated, modeled, smoothed, or extrapolated. Rounded rates are calculated from the published integer counts, and UNKNOWN remains part of the panel partition.
Put the weekly signal to work
For an SEO director, content operations lead, publisher, or retrieval team, this benchmark is context for a policy review. It cannot decide the organization's policy. A useful review names the crawler token, the paths in scope, the business purpose, the owner of the decision, and the date the rule should be checked again.
US Tech Automations can automate that review trail: capture an approved public policy, detect a later change, route it to the responsible owner, and preserve the decision beside the affected path. The workflow should keep policy evidence separate from crawl logs, citations, and model-output measurements, because those sources answer different questions.
The benchmark can also prevent overreaction. A competitor's root rule does not explain its full visibility strategy. A rising aggregate does not justify copying a block. Teams should inspect their own paths, products, and obligations before changing a directive. The value of the weekly series is a consistent outside reference, not an automatic instruction.
For organizations building an auditable handoff rather than another dashboard, the broader agentic workflow pattern shows how collection, review, approval, and evidence retention can fit together. The public index remains descriptive; the organization remains responsible for its policy.
Frequently Asked Questions
Q: Does the GPTBot classified-policy rate describe all top sites?
A: No. The 14.8% figure is the share of classified policies whose applicable root-path rule disallowed GPTBot. The full-panel share was 10.2%, and 30.8% of the panel was UNKNOWN. Transport-error missingness is systematic; classified-policy rates describe the answered subpopulation, while panel and UNKNOWN shares use the full frozen panel.
Q: Why not count UNKNOWN as not blocked?
A: Because the stored observation did not support that conclusion. Network failures, server failures, truncated bodies, HTML-shaped responses, missing bodies, and invalid text are evidence gaps. Turning those gaps into a permissive verdict would manufacture confidence.
Q: Does NOT_DISALLOWED_BY_TESTED_RULE mean a crawler has permission?
A: No. It describes only the tested robots rule for the named token at /. Robots instructions do not grant authorization, and another path or control may differ.
Q: Can I compare the prior headline directly with this edition's rate?
A: No. The old page used a different parser, question, and denominator. This in-place revision removes that headline rather than presenting a false trend. Future weekly updates can be compared when they use this same frozen method.
Q: Can the public download identify which domains changed?
A: No. The public JSON and CSV contain aggregate counts, definitions, dates, and integrity hashes. Domain identities, stored bodies, and per-domain verdicts remain outside the public artifact.
Q: What should happen if the source database or a hash cannot be checked?
A: The rebuild exits with UNKNOWN and writes no partial artifact. A missing source, schema mismatch, corrupt body, row-hash mismatch, batch-hash mismatch, or output drift cannot become a zero or a pass.
Source: US Tech Automations Research, Closing Web weekly read model, sealed observations for August 4 and August 11, 2026 over the frozen Tranco panel; aggregate root-path policy verdicts only.
Get this data as a daily feed
The numbers in this report come from a permit feed we monitor daily. Leave your email and we will follow up about a daily feed for your ZIPs and categories.
Prefer to talk first? Contact us.
Cite this report
US Tech Automations Research, 2026-08 edition. “AI Crawler Blocking by Bot: Weekly Closing Web Index.” https://ustechautomations.com/resources/blog/ai-crawler-blocking-trend-august-2026
Sealed snapshot sha256: 5ba00c28b989c6da9883d58a11770fbad7a5a435edb5f8c0b4ebd8d2729835e2
Machine-readable data: CSV · JSON · All research & methodology