US Tech Automations

Software and AI pages AI-crawler policy changes 59,000 rows held

What is and is not in the crawler feed

Every day up to 24 Aug 2026 we asked 59,000 sites for four files and saved what came back. This page says how many answered, how many said there is no file, how many refused us, and what we leave out before anything is counted.

Price
Not for sale
Built for
Publishers and SEO/AI teams watching who got blocked
Read
Every day, up to 24 Aug 2026
Newest sealed read
2026-08-30

Email us about the copies we holdNo card needed to ask. We reply with what we hold.

Newest sealed read: 2026-08-30. We read this source every day up to 24 Aug 2026, and collection has paused since then with no date set for it to start again. We hold 192 sealed runs going back to 2026-06-09.

What this page is59,000 rows held · newest sealed read 2026-08-30

  • We hold 29,765,400 dated rows in all, from 192 sealed runs between 9 Jun 2026 and 30 Aug 2026.
  • We asked every site for four files each day. On 30 Aug 2026, our last read, the number that had one: robots.txt 40,576, llms.txt 10,554, ai.txt 6,003, tdmrep.json 5,363.
  • Of the files we got back on 30 Aug 2026, 36,994 are a real robots.txt. Of the rest, 2,415 are a web page or a bot-block screen sent in its place and 1,105 are plain text with no User-agent line in them at all. We do not read either as a policy.
  • 10,225 of the real files give at least one AI crawler a User-agent line of its own. A further 74 mention one somewhere else in the file, in a comment or inside a path, which is not the same as naming it and is not counted here.
  • In the last 7 days, 13,075 robots.txt files changed at all. Most of those changes touch nothing an AI crawler cares about; 3,103 of them moved a named AI crawler.

Real rows out of our sealed copies

These are rows we read out of dated copies we keep ourselves. The live source shows today only, so once a row moves, what it said before is gone from the place you would go to look.

What the 59,000 sites did when we asked for robots.txt on our newest read. The indented lines break down the sites that never answered; none of those is a block and none is permission, it means we do not know. 30 Aug 2026
What the site didHow many sites
gave us the file (200)40,576
never answered at all11,036
— the address does not resolve3,753
— its certificate would not open3,223
— it never replied in time2,327
— it refused the connection1,161
— it cut the connection411
— the connection failed158
— httpexception1
— the address is not usable1
— unicodeerror1
said there is no such file (404)3,308
refused us (403)3,271
accepted but sent no file (202)271
told us to slow down (429)133
their server was unavailable (503)69
their server errored (500)40
called our request bad (400)35
answered with code 410 (410)29
answered with code 502 (502)29
answered with code 302 (302)19
The last 8 days we ran, and how many sites gave us a robots.txt each day. We have sealed 192 runs across 77 days since 9 Jun 2026. 23 Aug 2026 to 30 Aug 2026
DaySites that gave us a file
23 Aug 202661,450
24 Aug 202661,439
25 Aug 20261,728
26 Aug 202647,689
27 Aug 202648,188
28 Aug 202648,187
29 Aug 202648,152
30 Aug 202640,514
Crawler names that changed status somewhere in the last 7 days. Top 12 of 46 names that moved. 23 Aug 2026 to 30 Aug 2026
Crawler name, as sites write itChanges we caught
gptbot205
chatgpt-user204
perplexitybot202
claudebot194
google-extended187
ccbot178
meta-externalagent172
applebot-extended172
bytespider161
amazonbot153
oai-searchbot146
anthropic-ai105

What this page cannot tell you

  • Nothing on this page is newer than 30 Aug 2026. We read the 59,000 sites on one public list of the most-visited sites, and a site that was never on that list cannot appear here. That list was withdrawn on 24 Aug 2026 and collection has paused while it is replaced, so what you are reading is a dated archive with an end, not a running feed.
  • We can only show a change between two of our own reads. A site that changed its file and changed it back between two of our reads did something we never saw, and it is not on this page.
  • A crawler that is not named in the file is not the same as one that is allowed. When a crawler is not named, whatever the file says under User-agent: * applies to it instead. Us saying a name was dropped means only that: the name is gone.
  • A site with no robots.txt at all has not given permission and has not blocked anyone. It has said nothing. We count those separately and never read them as a yes or a no.
  • On our newest read, 3,271 sites refused us and 11,036 never answered. We do not know what their file said that day, and we do not guess.
  • Two kinds of change are left out rather than sold to you. 1,165 came from sites whose file swings between an old and a new version day to day, which is two of the site's servers disagreeing, not a change of mind. 11 touched a file too big for us to save whole, so we could not be sure we had read all of it.
  • What robots.txt says and what a crawler does are two different things. We report the file. We do not watch anyone's traffic, and we do not judge whether a site is right to block anyone.
  • 9,539 file changes in this window were left out because one of the two days answered with something we cannot read as a policy: a web page or a bot-block screen in 9,355 of them, and plain text with no User-agent line in the other 184.

See the file we hold

You do not have to take our word for what is in the file. Here are 12 rows of the real thing, carrying all 7 of its columns, cut out of the dated copies we sealed ourselves. Nothing in it is made up and nothing in it is tidied up.

  • Open the 12 rows as a CSVA plain spreadsheet file. It saves to your machine rather than painting itself into a browser tab, and it opens in Excel, Numbers or Google Sheets.
  • The same 12 rows as JSONThe same rows again, laid out for reading with code.

Nothing on this page is for sale. The sample file is free to read, and so is every row on the page above it.

These 12 rows are a slice of the file, not the whole of it. What we cannot show you here is how far back it goes: the file goes back further than these rows do.

Start the thread

No pay button on this one yet. Email operations@ustechautomations.com. We reply with what we hold for this one, and with what we do not hold, before you spend anything.

Email us about the copies we hold

Say that you want What is and is not in the crawler feed and we will tell you which weeks we hold for it and since when.

More from this feed