US Tech Automations

Software and AI pages AI-crawler policy changes 303 rows held

Sites that stopped naming an AI crawler in robots.txt

These 46 sites named an AI crawler in robots.txt one day and did not name it the next. That happened 303 times, counting each crawler separately, in the 7 days from 23 Aug 2026 to 30 Aug 2026. This is the one people read wrong: a name disappearing is not the same as a block coming off, because whatever the file says for every crawler at once then applies instead.

Price
Not for sale
Built for
Publishers and SEO/AI teams watching who got blocked
Read
Every day, up to 24 Aug 2026
Newest sealed read
2026-08-30

Email us about the copies we holdNo card needed to ask. We reply with what we hold.

Newest sealed read: 2026-08-30. We read this source every day up to 24 Aug 2026, and collection has paused since then with no date set for it to start again. We hold 192 sealed runs going back to 2026-06-09.

What this page is303 rows held · newest sealed read 2026-08-30

  • 303 changes across 46 named sites in the 7 days from 23 Aug 2026 to 30 Aug 2026.
  • The crawlers this happened to most: GPTBot (29), Google-Extended (25), CCBot (24), ClaudeBot (24), Bytespider (21).
  • The biggest single rewrite was taskade.com on 28 Aug 2026: 21 AI crawlers moved the same way in one edit.
  • 36 different crawler names are involved, all of them read out of the saved files rather than from a list we made up.

Real rows out of our sealed copies

These are rows we read out of dated copies we keep ourselves. The live source shows today only, so once a row moves, what it said before is gone from the place you would go to look.

One row per site, 12 sites shown of 46. The file you buy carries all 303 changes, with the site, the crawler, both readings and both dates. 23 Aug 2026 to 30 Aug 2026
SiteCrawlerWhat the file said beforeWhat it says nowBetween
msfn.orgAmazonbotblocked from the whole sitenot named in the file29 Aug 2026 → 30 Aug 2026
saleshandy.comFacebookBotblocked from 1 foldernot named in the file29 Aug 2026 → 30 Aug 2026
tcpdf.orgAiHitBotblocked from the whole sitenot named in the file29 Aug 2026 → 30 Aug 2026
nuvei.comGoogle-Extendedblocked from 13 foldersnot named in the file28 Aug 2026 → 29 Aug 2026
rssing.comanthropic-aiblocked from the whole sitenot named in the file28 Aug 2026 → 29 Aug 2026
turnkeylinux.orgmeta-externalagentblocked from 1 foldernot named in the file28 Aug 2026 → 29 Aug 2026
arynews.tvFacebookBotnamed, with 1 allow line and no blocknot named in the file27 Aug 2026 → 28 Aug 2026
beeradvocate.comGoogle-Extendednamed, with 1 allow line and no blocknot named in the file27 Aug 2026 → 28 Aug 2026
bitdegree.orgGPTBotblocked from the whole sitenot named in the file27 Aug 2026 → 28 Aug 2026
dxomark.comApplebot-Extendedblocked from the whole sitenot named in the file27 Aug 2026 → 28 Aug 2026
foreignaffairs.comAmazonbotblocked from the whole sitenot named in the file27 Aug 2026 → 28 Aug 2026
foreignaffairs.orgAmazonbotblocked from the whole sitenot named in the file27 Aug 2026 → 28 Aug 2026

What this page cannot tell you

  • Nothing on this page is newer than 30 Aug 2026. We read the 59,000 sites on one public list of the most-visited sites, and a site that was never on that list cannot appear here. That list was withdrawn on 24 Aug 2026 and collection has paused while it is replaced, so what you are reading is a dated archive with an end, not a running feed.
  • We can only show a change between two of our own reads. A site that changed its file and changed it back between two of our reads did something we never saw, and it is not on this page.
  • A crawler that is not named in the file is not the same as one that is allowed. When a crawler is not named, whatever the file says under User-agent: * applies to it instead. Us saying a name was dropped means only that: the name is gone.
  • A site with no robots.txt at all has not given permission and has not blocked anyone. It has said nothing. We count those separately and never read them as a yes or a no.
  • On our newest read, 3,271 sites refused us and 11,036 never answered. We do not know what their file said that day, and we do not guess.
  • Two kinds of change are left out rather than sold to you. 1,165 came from sites whose file swings between an old and a new version day to day, which is two of the site's servers disagreeing, not a change of mind. 11 touched a file too big for us to save whole, so we could not be sure we had read all of it.
  • What robots.txt says and what a crawler does are two different things. We report the file. We do not watch anyone's traffic, and we do not judge whether a site is right to block anyone.

See the file we hold

You do not have to take our word for what is in the file. Here are 12 rows of the real thing, carrying all 7 of its columns, cut out of the dated copies we sealed ourselves. Nothing in it is made up and nothing in it is tidied up.

  • Open the 12 rows as a CSVA plain spreadsheet file. It saves to your machine rather than painting itself into a browser tab, and it opens in Excel, Numbers or Google Sheets.
  • The same 12 rows as JSONThe same rows again, laid out for reading with code.

Nothing on this page is for sale. The sample file is free to read, and so is every row on the page above it.

These 12 rows are a slice of the file, not the whole of it. What we cannot show you here is how far back it goes: the file goes back further than these rows do.

Start the thread

No pay button on this one yet. Email operations@ustechautomations.com. We reply with what we hold for this one, and with what we do not hold, before you spend anything.

Email us about the copies we hold

Say that you want Sites that stopped naming an AI crawler and we will tell you which weeks we hold for it and since when.

More from this feed