Newest sealed read: 2026-08-30. We read this source every day up to 24 Aug 2026, and collection has paused since then with no date set for it to start again. We hold 192 sealed runs going back to 2026-06-09.
What this page is59,000 rows held · newest sealed read 2026-08-30
- We hold 29,765,400 dated rows in all, from 192 sealed runs between 9 Jun 2026 and 30 Aug 2026.
- We asked every site for four files each day. On 30 Aug 2026, our last read, the number that had one: robots.txt 40,576, llms.txt 10,554, ai.txt 6,003, tdmrep.json 5,363.
- Of the files we got back on 30 Aug 2026, 36,994 are a real robots.txt. Of the rest, 2,415 are a web page or a bot-block screen sent in its place and 1,105 are plain text with no User-agent line in them at all. We do not read either as a policy.
- 10,225 of the real files give at least one AI crawler a User-agent line of its own. A further 74 mention one somewhere else in the file, in a comment or inside a path, which is not the same as naming it and is not counted here.
- In the last 7 days, 13,075 robots.txt files changed at all. Most of those changes touch nothing an AI crawler cares about; 3,103 of them moved a named AI crawler.
Real rows out of our sealed copies
These are rows we read out of dated copies we keep ourselves. The live source shows today only, so once a row moves, what it said before is gone from the place you would go to look.
| What the site did | How many sites |
|---|---|
| gave us the file (200) | 40,576 |
| never answered at all | 11,036 |
| — the address does not resolve | 3,753 |
| — its certificate would not open | 3,223 |
| — it never replied in time | 2,327 |
| — it refused the connection | 1,161 |
| — it cut the connection | 411 |
| — the connection failed | 158 |
| — httpexception | 1 |
| — the address is not usable | 1 |
| — unicodeerror | 1 |
| said there is no such file (404) | 3,308 |
| refused us (403) | 3,271 |
| accepted but sent no file (202) | 271 |
| told us to slow down (429) | 133 |
| their server was unavailable (503) | 69 |
| their server errored (500) | 40 |
| called our request bad (400) | 35 |
| answered with code 410 (410) | 29 |
| answered with code 502 (502) | 29 |
| answered with code 302 (302) | 19 |
| Day | Sites that gave us a file |
|---|---|
| 23 Aug 2026 | 61,450 |
| 24 Aug 2026 | 61,439 |
| 25 Aug 2026 | 1,728 |
| 26 Aug 2026 | 47,689 |
| 27 Aug 2026 | 48,188 |
| 28 Aug 2026 | 48,187 |
| 29 Aug 2026 | 48,152 |
| 30 Aug 2026 | 40,514 |
| Crawler name, as sites write it | Changes we caught |
|---|---|
| gptbot | 205 |
| chatgpt-user | 204 |
| perplexitybot | 202 |
| claudebot | 194 |
| google-extended | 187 |
| ccbot | 178 |
| meta-externalagent | 172 |
| applebot-extended | 172 |
| bytespider | 161 |
| amazonbot | 153 |
| oai-searchbot | 146 |
| anthropic-ai | 105 |
What this page cannot tell you
- Nothing on this page is newer than 30 Aug 2026. We read the 59,000 sites on one public list of the most-visited sites, and a site that was never on that list cannot appear here. That list was withdrawn on 24 Aug 2026 and collection has paused while it is replaced, so what you are reading is a dated archive with an end, not a running feed.
- We can only show a change between two of our own reads. A site that changed its file and changed it back between two of our reads did something we never saw, and it is not on this page.
- A crawler that is not named in the file is not the same as one that is allowed. When a crawler is not named, whatever the file says under User-agent: * applies to it instead. Us saying a name was dropped means only that: the name is gone.
- A site with no robots.txt at all has not given permission and has not blocked anyone. It has said nothing. We count those separately and never read them as a yes or a no.
- On our newest read, 3,271 sites refused us and 11,036 never answered. We do not know what their file said that day, and we do not guess.
- Two kinds of change are left out rather than sold to you. 1,165 came from sites whose file swings between an old and a new version day to day, which is two of the site's servers disagreeing, not a change of mind. 11 touched a file too big for us to save whole, so we could not be sure we had read all of it.
- What robots.txt says and what a crawler does are two different things. We report the file. We do not watch anyone's traffic, and we do not judge whether a site is right to block anyone.
- 9,539 file changes in this window were left out because one of the two days answered with something we cannot read as a policy: a web page or a bot-block screen in 9,355 of them, and plain text with no User-agent line in the other 184.
See the file we hold
You do not have to take our word for what is in the file. Here are 12 rows of the real thing, carrying all 7 of its columns, cut out of the dated copies we sealed ourselves. Nothing in it is made up and nothing in it is tidied up.
- Open the 12 rows as a CSVA plain spreadsheet file. It saves to your machine rather than painting itself into a browser tab, and it opens in Excel, Numbers or Google Sheets.
- The same 12 rows as JSONThe same rows again, laid out for reading with code.
Nothing on this page is for sale. The sample file is free to read, and so is every row on the page above it.
These 12 rows are a slice of the file, not the whole of it. What we cannot show you here is how far back it goes: the file goes back further than these rows do.
Start the thread
No pay button on this one yet. Email operations@ustechautomations.com. We reply with what we hold for this one, and with what we do not hold, before you spend anything.
Email us about the copies we holdSay that you want What is and is not in the crawler feed and we will tell you which weeks we hold for it and since when.
More from this feed
- Up one level: AI-crawler policy changesThe whole feed, its price, and how the file arrives.