27,210 AI-Crawler Requests: Who Reads Our Pages
Every AI company that reads the open web leaves a trace in the same place: a web server's access log. Nobody outside a company can see that log, which is exactly why almost nobody publishes this kind of count. The AI-Crawler Census is our own answer to one question — of the requests that hit our own pages, which named AI companies sent them, and how many. On August 25, 2026, the crawler-census clock sealed 24 days of its own server access logs and published the count without alteration.
Across three tracked web surfaces, 27,210 requests arrived from AI-company crawlers between August 1, 2026 and August 24, 2026. This is a count of requests we received. It is not, and must never be read as, a measure of what any AI product then showed, cited, or ranked as a result.
A crawler request is one automated fetch of one page. It is not a reader, a visitor, or a person.
Key Takeaways
10 named AI companies sent 27,210 requests across 3 tracked web surfaces in 24 days. Meta leads at a 26% share, with Anthropic close behind at 20.8%.
12 distinct crawler bots were observed, each one attributed to the company that operates it, from Meta-ExternalAgent down to DuckAssistBot.
One of the three surfaces — the main site — is under-measured this edition: its one available day was read under a 5,000-row cap, and 23 of its 24 days are absent from these totals entirely, not zero.
Rankings pages drew the largest share of crawler attention at 42.2%, ahead of every other kind of page we publish.
This edition draws a hard line at August 25, 2026: main-site counts from before that date are floors, not counts, and trending across that boundary would measure our own logging fix, not crawler behavior.
Where the Requests Landed
US Tech Automations runs three separate web surfaces, and the census tracks each one on its own terms rather than blending them into a single number that would hide how differently each surface is measured.
| Surface | Days available | Days unavailable | Requests | Share | Distinct paths | Count basis |
|---|---|---|---|---|---|---|
| permits.ustechautomations.com | 24 | 0 | 25,937 | 95.3% | 2,504 | a complete count: the whole day log was read, with no row cap |
| ustechautomations.com (main site) | 1 | 23 | 1,161 | 4.3% | 1,065 | a floor, not a count: 1 of 1 available days were read under a 5,000-row log cap |
| agents.ustechautomations.com | 24 | 0 | 112 | 0.4% | 18 | a complete count: the whole day log was read, with no row cap |
25,937 of the edition's requests, a 95.3% share, landed on the permits surface alone. That surface has the cleanest measurement in the edition: 24 of 24 days available, no cap, a genuine complete count. The main site is the opposite case — 1 day available against 23 unavailable, and even that single day was read under a cap, so its 1,161 requests and 4.3% share are a floor on what actually happened, not the real total.
The agent-market surface is the smallest by volume, at 112 requests and a 0.4% share, but it shares the permits surface's clean measurement: 24 of 24 days, no cap, no gap.
Which AI Companies Showed Up
Ten named companies sent crawler traffic to our surfaces this edition. The table rolls every crawler bot up to the company that operates it.
| AI company | Requests | Share |
|---|---|---|
| Meta | 7,076 | 26% |
| Anthropic | 5,670 | 20.8% |
| OpenAI | 5,067 | 18.6% |
| Amazon | 3,608 | 13.3% |
| ByteDance | 3,009 | 11.1% |
| Perplexity | 1,787 | 6.6% |
| 807 | 3% | |
| Common Crawl | 104 | 0.4% |
| You.com | 54 | 0.2% |
| DuckDuckGo | 28 | 0.1% |
Meta sent 7,076 requests this edition, a 26% share — the largest of any named company. Anthropic and OpenAI follow closely, at 20.8% and 18.6% respectively, so the top three companies alone account for most of the traffic in the edition. The next tier — Amazon at 13.3% and ByteDance at 11.1% — still moves real volume, while Google, Common Crawl, You.com, and DuckDuckGo each trail well behind Perplexity's 1,787.
For site owners deciding which crawlers to allow, our robots.txt data on who blocks Meta-ExternalAgent shows how common it already is to shut that traffic out — a decision this census can't make for you, only inform.
Crawler by Crawler
The company rollup hides real variation underneath it: some companies run one crawler, others run several with different jobs.
| Crawler | Company | Requests | Share |
|---|---|---|---|
| Meta-ExternalAgent | Meta | 7,076 | 26% |
| Anthropic ClaudeBot | Anthropic | 5,670 | 20.8% |
| OpenAI ChatGPT-User | OpenAI | 4,027 | 14.8% |
| Amazonbot | Amazon | 3,608 | 13.3% |
| ByteDance Bytespider | ByteDance | 3,009 | 11.1% |
| PerplexityBot | Perplexity | 1,787 | 6.6% |
| GoogleOther | 807 | 3% | |
| OpenAI SearchBot | OpenAI | 680 | 2.5% |
| OpenAI GPTBot | OpenAI | 360 | 1.3% |
| Common Crawl CCBot | Common Crawl | 104 | 0.4% |
| You.com YouBot | You.com | 54 | 0.2% |
| DuckAssistBot | DuckDuckGo | 28 | 0.1% |
OpenAI is the clearest example of one company running several distinct jobs: ChatGPT-User at 14.8%, SearchBot at 2.5%, and GPTBot at 1.3% are three separate bots with three separate purposes, and together they add up to more than any single-bot company's share. A site that blocks "OpenAI" in the abstract has to decide whether it means all three, and our writeup on who blocks PerplexityBot covers the same one-bot-at-a-time decision from the blocking side.
What Kind of Page Got Read
Crawler traffic did not land evenly across the site. The census groups every requested path into one of several page families.
| Page family | Requests | Share | Distinct paths |
|---|---|---|---|
| rankings pages | 11,483 | 42.2% | 1,372 |
| other pages | 7,244 | 26.6% | 787 |
| open data feeds | 4,468 | 16.4% | 116 |
| product and offer pages | 2,114 | 7.8% | 195 |
| blog posts | 1,073 | 3.9% | 1,012 |
| robots.txt, llms.txt and sitemaps | 722 | 2.7% | 86 |
| the home page | 82 | 0.3% | 2 |
| agent-protocol endpoints | 18 | 0.1% | 7 |
| vulnerability-probe paths | 3 | 0% | 3 |
| research reports | 3 | 0% | 3 |
Rankings pages alone absorbed 42.2% of every crawler request in the edition. Open data feeds are the second-most efficient family to watch: 4,468 requests concentrated on just 116 distinct paths, versus blog posts, where 1,073 requests spread across 1,012 distinct paths — nearly one request per page.
That contrast says something about how crawlers browse: feeds get hit hard and repeatedly, while the long tail of blog content gets a lighter, broader pass. The vulnerability-probe row is worth naming rather than hiding: 3 requests were classified as probes for paths like /.env, reported under their own family instead of being folded into legitimate crawler traffic.
Some of that lighter blog-post pass is by design rather than accident — our Google-Extended blocking data shows how many sites choose to shut a given crawler out of certain sections entirely, which thins the traffic those crawlers can even reach.
The Floor, the Gap, and the Boundary We Won't Cross
This edition's most important number is the one it refuses to state as a fact: how many crawler requests actually hit the main site before August 25, 2026.
Before August 25, 2026, main-site counts are a floor, not a count. The real number is that or higher, and it is unknown by how much.
The main site's log reader was capped at 5,000 rows per day, and for 23 of the 24 days in this window, the main site's log was unavailable entirely — those days are absent from every total above, not zero. A day the log reader could not reach is a day we did not measure, and reporting it as zero would be publishing something false. The one main-site day we could read, we read under that same 5,000-row cap, so even its 1,161 requests are a floor on the true number, not the number itself.
Because of that fix, we draw one hard line: nothing here trends across August 25, 2026. Main-site counts before that date are truncated floors; counts from that day forward are complete. Comparing the two would measure our own instrumentation getting fixed, not any change in how AI companies actually crawl us — and that is not a comparison this report will make, this edition or any future one.
What Counts as a Request
A request, in this census, is one logged HTTP fetch of one path on one of our own three surfaces, attributed to a named crawler by its user-agent string. A crawler name is the specific bot identity a company operates — Meta-ExternalAgent, ClaudeBot, GPTBot, and so on — rolled up to the company behind it for the company-level table above.
What this census excludes matters as much as what it includes. It does not exclude our own uptime probes on the permits and agent-market surfaces, which are filtered out before the count starts; the main-site surface has no such filter. And critically, a logged request tells us nothing about what happened after the fetch — whether the page was used to answer a question, cited in an output, or shown to anyone at all. We count the fetch. We do not, and cannot, count what came next.
Put the Crawler Data to Work
A raw crawler count is a curiosity on its own. It becomes useful the moment a content or SEO lead can compare it against what those same pages actually earn in search, and decide where to spend the next month of publishing effort. A page family drawing 42.2% of crawler attention but a thin share of clicks is a different problem than a page family nobody reads at all, and the two calls for different fixes — one is a relevance problem, the other is a discovery problem.
US Tech Automations builds the workflows that make that comparison a standing job instead of a one-off pull: an automated pipeline that reads server logs and search-performance data side by side, on a schedule, so a content strategy lead does not have to reconcile two different exports by hand every month. The agentic workflows platform is where that kind of recurring, cross-source reporting gets built once and then runs on its own.
Frequently Asked Questions
Q: How many AI companies were reading our pages this edition?
A: 10 named companies sent crawler traffic across 12 distinct bots this edition, covering 27,210 requests over 24 days. Meta led at a 26% share, followed by Anthropic at 20.8% and OpenAI at 18.6%.
Q: Does this count show what ChatGPT, Claude, or any other AI product actually told a user?
A: No. This census counts requests we received on our own servers. It cannot see, and does not claim to know, whether any of those requests led to a page being cited, summarized, or shown to anyone. That is a different kind of data, and we do not have it.
Q: Why is the main site's crawler count lower than the permits site's?
A: Largely because it is under-measured, not because crawlers ignore it. The permits surface has a complete 24-day log with no cap. The main site has only 1 available day out of 24, and even that day was read under a 5,000-row cap, so its 1,161 requests and 4.3% share are a floor, not the real total.
Q: Can I compare this edition's main-site numbers to next month's?
A: Not across the August 25, 2026 boundary. Main-site logging was fixed on that date, so counts before it are floors and counts after it are complete. A rise across that line would reflect the fix, not a real change in crawler behavior.
Q: Are all 27,210 requests from real AI companies, or is some of it noise?
A: The total includes 3 requests classified as vulnerability probes — automated scans for paths like /.env — reported honestly under their own page family rather than mixed into legitimate crawler traffic. The rest are attributed to one of the 12 named crawler bots in the table above.
Method and Provenance
Every figure above is counted directly from sealed daily census files; nothing is estimated, modeled, or extrapolated. Where a day was not collected it is reported as missing, never as zero, and where a day was read under a row cap it is reported as a floor, never as a count. This count covers requests to our own web surfaces from named AI-company crawlers, taken from our own server access logs — it counts requests received, not outcomes downstream of them.
Source: US Tech Automations Research — AI-Crawler Census, sealed edition 2026-08, from our own server access logs across permits.ustechautomations.com, ustechautomations.com, and agents.ustechautomations.com.
Get this data as a daily feed
The numbers in this report come from a permit feed we monitor daily. Leave your email and we will follow up about a daily feed for your ZIPs and categories.
Prefer to talk first? Contact us.
Cite this report
US Tech Automations Research, 2026-08 edition. “27,210 AI-Crawler Requests: Who Reads Our Pages.” https://ustechautomations.com/resources/blog/which-ai-companies-read-our-pages
Sealed snapshot sha256: efb0feca4c66af91
Machine-readable data: CSV · JSON · All research & methodology
About the Author

Helping businesses leverage automation for operational efficiency.