8 of 37 Named Business Sites Block AI Crawlers: Benchmark
US Tech Automations Research audited a fixed cohort of 40 named Accounting, HR, Marketing, and Productivity sites on July 20, 2026. Of the 37 sites returning a parseable robots.txt, 8 publish an exact tracked user-agent group with Disallow: /. Separately, llms.txt was observed on 18 of all 40 named sites.
Scope: A point-in-time audit of the named Accounting, HR, Marketing, and Productivity sites in the curated Closing Web cohort. The cohort includes vendors, trade publications, and associations; it is not a representative sample of software vendors, business sites, or the web.
A block means an exact tracked user-agent group with Disallow: / in the site public robots.txt. Blocking rates use only sites with a parseable robots.txt; llms.txt observations use all named sites. These files express or publish instructions, not proof of enforcement, crawler compliance, crawl traffic, citations, rankings, search inclusion, or visibility impact. llms.txt is a proposal, and its presence does not prove an LLM consumes it or that it improves GEO.
Key Takeaways
AI blocking: 8 of 37 parseable policies, or 21.6%. This denominator excludes sites whose robots.txt was not parseable.
llms.txt observed: 18 of 40 named sites, or 45%. That denominator includes the entire named cohort, regardless of robots.txt status.
Accounting records 4 of 8 parseable policies blocking a tracked token, while HR records 2 of 9. Marketing and Productivity each record 1 of 10.
Bytespider appears in 5 of 37 parseable policies, while PerplexityBot appears in 0 of 37 under the exact-token, sitewide-block definition.
The cohort tracks 21 exact user-agent tokens, but the headline bot table reports the specified 9 tokens. A policy for one token should not be generalized to a provider's other search, training, or user-triggered agents.
Blocking and llms.txt use different denominators: 37 parseable robots.txt policies versus all 40 named sites.
The Direct Answer by Category
The cohort-level answer hides meaningful category differences, but those differences still describe only these named sites. Accounting includes vendors, publications, and an association; the same mixed-site principle applies elsewhere. A category percentage is therefore a cohort descriptor, not an industry estimate.
| Category | Named sites | Parseable robots.txt | Any exact-token block | Block rate | llms.txt observed | llms.txt rate |
|---|---|---|---|---|---|---|
| Accounting | 10 | 8 | 4 of 8 | 50% | 2 of 10 | 20% |
| HR | 10 | 9 | 2 of 9 | 22.2% | 3 of 10 | 30% |
| Marketing | 10 | 10 | 1 of 10 | 10% | 5 of 10 | 50% |
| Productivity | 10 | 10 | 1 of 10 | 10% | 8 of 10 | 80% |
The dedicated Accounting crawler audit and HR crawler audit provide category-level context.
The Marketing crawler audit and Productivity crawler audit do the same for the two categories with 10 parseable policies.
This benchmark is intentionally not another category fanout. It holds one named cohort together, exposes every row, and separates robots.txt blocking from llms.txt observation. That makes the result auditable without suggesting that four mixed groups describe the broader business web.
Exact-Token Blocking, Not Provider-Level Guessing
The audit counts a block only when the robots.txt contains a group naming the tracked token exactly and that group contains Disallow: /. A partial-path rule is not the sitewide condition measured here. A generic group is not silently reassigned to a named token. A differently capitalized token remains distinct when the snapshot tracks it distinctly.
| Tracked token | Exact sitewide blocks | Share of 37 parseable policies |
|---|---|---|
| GPTBot | 3 of 37 | 8.1% |
| ClaudeBot | 2 of 37 | 5.4% |
| Google-Extended | 2 of 37 | 5.4% |
| CCBot | 4 of 37 | 10.8% |
| PerplexityBot | 0 of 37 | 0% |
| Bytespider | 5 of 37 | 13.5% |
| Meta-ExternalAgent | 2 of 37 | 5.4% |
| Amazonbot | 4 of 37 | 10.8% |
| Applebot-Extended | 2 of 37 | 5.4% |
Bytespider blocks: 5 of 37 parseable policies, or 13.5%. That is an exact-token result. It does not prove those sites successfully denied requests, that Bytespider attempted to crawl them, or that another ByteDance user agent received the same instruction.
The zero for PerplexityBot is equally narrow. It means 0 of 37 parseable policies met this report's exact PerplexityBot plus Disallow: / condition. It does not mean every Perplexity request was allowed, that no other access control existed, or that every site appeared in a Perplexity answer.
Search, Training, and User-Triggered Agents Are Different
Official crawler documentation makes purpose separation essential. OpenAI's crawler documentation describes OAI-SearchBot as a search crawler, GPTBot as a crawler for content that may be used in model training, and ChatGPT-User as user-triggered rather than an automatic web crawler. A GPTBot rule therefore should not be rewritten as a blanket statement about ChatGPT search.
Anthropic's crawler documentation likewise distinguishes ClaudeBot for potential model-training content, Claude-SearchBot for search, and Claude-User for user-initiated retrieval. The dataset tracks those exact labels separately; the headline table does not collapse them into a single Anthropic rate.
Google's official crawler reference describes Google-Extended as a standalone robots.txt product token used to manage model training and grounding uses. Google also says that token does not affect inclusion or ranking in Google Search. A Google-Extended block is not evidence of a conventional Google Search block.
| Purpose distinction | Example tracked tokens | What a policy row can support | What it cannot support |
|---|---|---|---|
| Model-development or training control | GPTBot, ClaudeBot, Google-Extended | The exact published instruction for that token | A claim about every product from the same provider |
| Search-oriented crawling | OAI-SearchBot, Claude-SearchBot | The exact published instruction for that search token | A claim about training-token policy |
| User-triggered retrieval | ChatGPT-User, Claude-User, Perplexity-User, MistralAI-User | The exact published instruction for that user token | Proof that robots.txt was enforced on a request |
| Other tracked crawlers | CCBot, Bytespider, Meta-ExternalAgent, Amazonbot, Applebot-Extended | The exact sitewide block condition in the snapshot | Crawl volume, compliance, visibility, or citation impact |
A block is a published instruction for one exact token, not proof of enforcement and not a provider-wide visibility verdict.
What llms.txt Observation Does and Does Not Mean
The llms.txt project describes the file as a proposal for presenting concise, LLM-friendly site information and links. The proposal does not establish a universal consumption requirement. This benchmark therefore records only whether llms.txt was observed at the checked location.
That observation is independent of robots.txt parseability. A site can have an observed llms.txt and publish no exact sitewide block for the headline tokens. Another can publish blocking instructions without an observed llms.txt. The files serve different purposes, and neither observation is a performance metric.
llms.txt observation rate: 45% across all 40 named sites. The rate does not measure file quality, freshness, completeness, retrieval, citations, GEO lift, rankings, or traffic. It also does not mean the unobserved sites intentionally rejected the proposal.
This separation matters for implementation. A web team choosing to publish llms.txt should do so because the file offers a useful machine-readable guide to authoritative content, not because this benchmark proves a ranking benefit. A team editing robots.txt should choose policy by crawler purpose and business intent, not by copying a peer's row.
Named-Site Audit Appendix
The appendix is the audit trail behind the aggregates. “None” means no tracked exact token met the report's sitewide block condition in a parseable robots.txt. “Not assessed” means robots.txt was not parseable, so the site is excluded from the blocking denominator. llms.txt remains assessed against the all-site denominator.
| Category | Domain | robots.txt status | Exact tracked tokens blocked sitewide | llms.txt |
|---|---|---|---|---|
| Accounting | accountingtoday.com | Parseable | Amazonbot | Not observed |
| Accounting | aicpa-cima.com | Parseable | CCBot, Bytespider | Not observed |
| Accounting | cpapracticeadvisor.com | Parseable | None | Not observed |
| Accounting | freshbooks.com | Parseable | Bytespider | Not observed |
| Accounting | hrblock.com | Not parseable | Not assessed | Not observed |
| Accounting | intuit.com | Parseable | None | Not observed |
| Accounting | journalofaccountancy.com | Parseable | CCBot, Bytespider | Not observed |
| Accounting | taxfoundation.org | Parseable | None | Observed |
| Accounting | waveapps.com | Not parseable | Not assessed | Not observed |
| Accounting | xero.com | Parseable | None | Observed |
| HR | adp.com | Parseable | None | Observed |
| HR | bamboohr.com | Parseable | None | Observed |
| HR | greenhouse.io | Parseable | None | Not observed |
| HR | gusto.com | Not parseable | Not assessed | Not observed |
| HR | hrdive.com | Parseable | None | Not observed |
| HR | hrexecutive.com | Parseable | GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, Meta-ExternalAgent, meta-externalagent, Amazonbot, Applebot-Extended | Observed |
| HR | lever.co | Parseable | None | Not observed |
| HR | shrm.org | Parseable | None | Not observed |
| HR | tlnt.com | Parseable | GPTBot | Not observed |
| HR | workday.com | Parseable | None | Not observed |
| Marketing | adweek.com | Parseable | GPTBot, ClaudeBot, anthropic-ai, Claude-Web, Google-Extended, CCBot, Bytespider, Meta-ExternalAgent, meta-externalagent, FacebookBot, Amazonbot, Applebot-Extended, cohere-ai | Not observed |
| Marketing | ahrefs.com | Parseable | None | Not observed |
| Marketing | buffer.com | Parseable | None | Not observed |
| Marketing | hootsuite.com | Parseable | None | Observed |
| Marketing | hubspot.com | Parseable | None | Observed |
| Marketing | mailchimp.com | Parseable | None | Observed |
| Marketing | marketingdive.com | Parseable | None | Not observed |
| Marketing | moz.com | Parseable | None | Not observed |
| Marketing | semrush.com | Parseable | None | Observed |
| Marketing | sproutsocial.com | Parseable | None | Observed |
| Productivity | airtable.com | Parseable | None | Not observed |
| Productivity | asana.com | Parseable | None | Observed |
| Productivity | clickup.com | Parseable | None | Observed |
| Productivity | evernote.com | Parseable | None | Observed |
| Productivity | monday.com | Parseable | None | Observed |
| Productivity | notion.so | Parseable | Amazonbot | Observed |
| Productivity | slack.com | Parseable | None | Observed |
| Productivity | todoist.com | Parseable | None | Not observed |
| Productivity | trello.com | Parseable | None | Observed |
| Productivity | zapier.com | Parseable | None | Observed |
The table shows why a named appendix is more useful than a bare rate. Accounting's 4 of 8 comes from different token combinations across accountingtoday.com, aicpa-cima.com, freshbooks.com, and journalofaccountancy.com. HR's 2 of 9 comes from hrexecutive.com and tlnt.com, with very different breadth. The same category aggregate can conceal distinct policy choices.
Marketing and Productivity also demonstrate why site type should remain visible. The report does not infer why adweek.com or notion.so publishes its particular rules. Business rationale, contract terms, server enforcement, authenticated content, and request logs are outside the sealed snapshot.
Methodology and Reproducibility
The source is the sealed curated Closing Web snapshot for July 20, 2026. US Tech Automations Research selected the fixed Accounting, HR, Marketing, and Productivity cohort, read each public robots.txt result, evaluated the 21 tracked exact tokens, and checked the llms.txt observation stored in the same snapshot.
For blocking, the denominator is 37 sites with a parseable robots.txt. A site counts for bot X only when an exact user-agent group names X and includes Disallow: /. The report does not broaden the match to provider families, partial paths, inferred aliases, or untracked controls.
For llms.txt, the denominator is all 40 named sites. “Observed” means the file was observed in the sealed check. It is not a quality review and does not prove downstream use. Because robots.txt and llms.txt answer different questions, their rates are never merged into a score.
Every count, percentage, domain, status, and token in the tables comes from the sealed snapshot; nothing is estimated, modeled, or extrapolated. The edition is cross-sectional only. It does not claim a trend from earlier category pages, even where the same domain appears in prior research.
Put the AI-Access Audit to Work
A marketing operations lead can use this benchmark as a schema for its own domain list: record the exact token, purpose, directive, scope, timestamp, and owner decision. A content lead can separately maintain llms.txt as an editorial map. Neither team should use a cohort percentage as an instruction to copy another site's policy.
The Productivity crawler audit shows how a named category report keeps site rows attached to its aggregate. The practical next step is a repeatable diff: fetch approved domains, parse exact tokens, preserve the raw response, compare it with the last sealed state, and route only meaningful changes to the policy owner.
US Tech Automations can orchestrate that process by scheduling fetches, storing policy versions, classifying token-level changes, and routing an approval task to an SEO, content, legal, or platform owner. It does not promise that a robots.txt or llms.txt change will improve rankings, citations, inclusion, or GEO.
The agentic workflows platform is the appropriate fit when a team needs monitored domain lists, durable diffs, approval history, and cross-system alerts. US Tech Automations is unnecessary when a webmaster can review a small domain set manually and the organization has no recurring policy-governance need.
Frequently Asked Questions
Q: What does “8 of 37 sites block AI crawlers” mean?
A: It means 8 of the 37 named sites with a parseable robots.txt publish at least one exact tracked user-agent group with Disallow: /. It does not mean 8 of every 37 business sites on the web block AI crawlers.
Q: Why does llms.txt use 40 sites instead of 37?
A: The llms.txt check is independent of robots.txt parsing, so its denominator is the full named cohort of 40. Blocking percentages exclude the sites whose robots.txt was not parseable; llms.txt observations do not.
Q: Does a robots.txt block prove that a crawler complied?
A: No. The audit measures published instructions. It does not inspect request logs, enforcement layers, crawler behavior, traffic, indexing, rankings, citations, or visibility impact.
Q: Does an observed llms.txt improve GEO or LLM citations?
A: This dataset cannot support that claim. llms.txt is a proposal, and observation does not prove that an LLM consumed the file or that it changed any ranking, answer, citation, or traffic outcome.
Q: Can GPTBot blocking be treated as a ChatGPT Search opt-out?
A: No. OpenAI documents GPTBot and OAI-SearchBot as separate tokens with different purposes. Policy analysis should preserve the exact token instead of turning one directive into a provider-wide conclusion.
Source: US Tech Automations Research — derived from the sealed curated Closing Web snapshot, exact user-agent plus Disallow: / methodology, edition 2026-07.
Get this data as a daily feed
The numbers in this report come from a permit feed we monitor daily. Leave your email and we will follow up about a daily feed for your ZIPs and categories.
Prefer to talk first? Contact us.
Cite this report
US Tech Automations Research, 2026-07 edition. “8 of 37 Named Business Sites Block AI Crawlers: Benchmark.” https://ustechautomations.com/resources/blog/business-sites-ai-crawler-access-benchmark-2026
Sealed snapshot sha256: 2914487c16f51e28b32bc6b3ed5aab9caef2c4b8796d91f718553d5d6e856d3c
Machine-readable data: CSV · JSON · All research & methodology
About the Author

Helping businesses leverage automation for operational efficiency.