US Tech Automations

Software and AI pages Daily seals, paused 24 Aug 2026 Sample ready

AI-crawler policy changes

Sites rewrite robots.txt with no notice and keep no history. We saved a copy of every file every day up to 24 Aug 2026, so you get the day a site let a crawler in or shut one out. Collection has paused since then; what we hold is dated and complete to that day.

Price
Not for sale
Built for
Publishers and SEO/AI teams watching who got blocked
Cadence
Daily copies up to 24 Aug 2026, paused since; the archive to that day is still available
Public sample
Named sites on this page

Email us about the copies we holdNo card needed to ask. We reply with what we hold.

Collection has pausedLast sealed copy 30 Aug 2026

Collection has paused. The list of sites this feed reads was withdrawn on 24 Aug 2026 and is being replaced, so nothing new has been collected since that day. This is not a feed running late and it is not catching up. We have set no date for it starting again, and we are not going to guess at one.

What we already hold is untouched and still available: 29,765,400 dated rows covering 134,179 sites over 77 days, from 9 Jun 2026 to 30 Aug 2026. Every row and every number on this page and its child pages is read out of that archive, so nothing below has changed and nothing below is going to move.

Public sample12 named sites · 320 sites changed their answer in these 8 reads

A site rewrites robots.txt whenever it likes and keeps no history of it. The file you fetch today is the only one there is, so the rule that applied to your crawler last week is not recoverable from the site. We saved a copy of each file every day up to 24 Aug 2026.

One row per site, and every kind of move we hold. 12 sites shown of 320. The file behind this page holds all 3,103 changes; what you buy is the part of it for the sites you name. 23 Aug 2026 to 30 Aug 2026
SiteCrawlerWhat the file said beforeWhat it says nowBetween
pourquoidocteur.frChatGPT-Userblocked from the whole sitenamed, with 1 allow line and no block27 Aug 2026 → 28 Aug 2026
ediblecommunities.comAmazonbotnot named in the fileblocked from the whole site29 Aug 2026 → 30 Aug 2026
msfn.orgAmazonbotblocked from the whole sitenot named in the file29 Aug 2026 → 30 Aug 2026
lazada.co.thChatGPT-Usernot named in the filenamed, with 1 allow line and no block29 Aug 2026 → 30 Aug 2026
akniga.orgClaude-SearchBotblocked from 17 foldersblocked from 15 folders29 Aug 2026 → 30 Aug 2026
fastcoexist.comMeta-ExternalAgentblocked from the whole siteblocked from 1 folder26 Aug 2026 → 27 Aug 2026
tbsradio.jpBytespidernot named in the fileblocked from the whole site29 Aug 2026 → 30 Aug 2026
saleshandy.comFacebookBotblocked from 1 foldernot named in the file29 Aug 2026 → 30 Aug 2026
audible.caGoogleOthernot named in the fileblocked from 8 folders28 Aug 2026 → 29 Aug 2026
betterimpact.comGPTBotblocked from 5 foldersblocked from 4 folders29 Aug 2026 → 30 Aug 2026
fastcompany.comMeta-ExternalAgentblocked from the whole siteblocked from 1 folder26 Aug 2026 → 27 Aug 2026
audioholics.comAmazonbotnot named in the fileblocked from the whole site28 Aug 2026 → 29 Aug 2026

Of the 12 sites shown, 3 let a crawler back in, 3 shut one out of the whole site, 2 stopped naming one at all, 2 named one for the first time and 2 rewrote the rules under one. A name disappearing is not the same as a block coming off, and that is the one people most often get wrong: when a crawler is not named, whatever the file says under User-agent: * applies to it instead.

The biggest single rewrite in this window35 crawlers moved in one edit

pcmag.com is the largest move we caught between two reads: on 28 Aug 2026 its file moved 35 named AI crawlers at once. One edit, one day, and nothing on the site to say it happened.

Across the whole window the crawlers this happened to most were GPTBot (205), ChatGPT-User (204), PerplexityBot (202), ClaudeBot (194), Google-Extended (187). Those names are read out of the saved files, spelled the way the sites themselves spell them, not matched against a list we wrote.

The size of the panel, and the honest count

We ask 59,000 sites for their robots.txt, and we hold a dated copy of the panel from 77 of the 83 days since 9 Jun 2026. 6 days are missing — a 2-day stretch from 11 Jun to 12 Jun, 19 Jun, 4 Jul, 15 Jul and 8 Aug 2026. We hold no copy from those days and no record of a run on them either. We are not going to invent a reason for that, and nothing on this page counts those days as read. On our newest read, 30 Aug 2026, 40,514 of them gave us a file we could save, 36,994 of those were a real robots.txt and not one of the 2,415 web pages or 1,105 plain files with no User-agent line sent in their place, and 10,225 of the real ones give at least one AI crawler a User-agent line of its own. Across the 8 reads this page counts, 23 Aug 2026 to 30 Aug 2026, 320 sites changed what their file says about a named AI crawler, across 3,103 crawler-level changes.

Every number in that paragraph was counted out of the stored files themselves when this page was built, and each one is shown again, broken down, on what is and is not in this feed. Ask us and we will show the working.

1,165 file changes on 406 sites were thrown out of this count on purpose. Those files swing between an old and a new version day to day, which is two of the site’s servers answering differently, not the site changing its mind. Counting them would inflate the number and every one of them would be a false alarm in your inbox.

9,539 more file changes were left out because one of the two days answered with something we cannot read as a policy: a web page or a bot-block screen in 9,355 of them, and plain text with no User-agent line in the other 184. Neither is a policy and we do not read either as one.

A very big file is saved up to a limit and no further. The limit is 512 KB. 114,492 of the 1,311,434 copies we hold are cut that way, and 11 changes in this window touched one, so we left those out rather than guess at the part we never read.

272 sites answered us on 30 Aug 2026 and left us nothing to keep, and 3,271 refused us outright. We do not know what any of their files said that day, and we do not guess.

Doing this yourself

You can fetch any site’s robots.txt yourself in a second. What you cannot do is fetch yesterday’s. Nothing about that file is versioned, announced or archived.

To watch a panel this size you would have to fetch 59,000 files every day, store every copy, and then work out which differences are a real policy change and which are two of the site’s servers disagreeing. That last part is most of the work: it is the 1,165 file changes above that we throw away.

What you get

  • Every site in your list that changed its answerSite, crawler, what it said before, what it says now, and both dates.
  • Sites that flap between two versions are held back, not sentYou get changes, not noise.
  • The saved copy behind any row, on requestSo you can read the whole rule, not just the part quoted on the page.
  • Cancel any month by emailNo account to close, no notice period.

How it works

  1. You email us the sites you care about, or ask for the whole panel.
  2. We tell you which of them we already hold and since when, then send a checkout link in that thread.
  3. A person emails you the changes file, and names anything we could not collect.

See the file we hold

You do not have to take our word for what is in the file. Here are 12 rows of the real thing, carrying all 7 of its columns, cut out of the dated copies we sealed ourselves. Nothing in it is made up and nothing in it is tidied up.

  • Open the 12 rows as a CSVA plain spreadsheet file. It saves to your machine rather than painting itself into a browser tab, and it opens in Excel, Numbers or Google Sheets.
  • The same 12 rows as JSONThe same rows again, laid out for reading with code.

Nothing on this page is for sale. The sample file is free to read, and so is every row on the page above it.

These 12 rows are a slice of the file, not the whole of it. What we cannot show you here is how far back it goes: the file goes back further than these rows do.

Start the thread

No pay button on this one yet. Email operations@ustechautomations.com. Send the list of sites you follow. We reply with which of them we hold and for which days. Nothing on this page is for sale.

Email us about the copies we hold

We will tell you which of your sites we already hold, and since when. There is nothing to pay for: this feed is not for sale.

Every page in this feed

What is and is not in the crawler feed