Counted from the archive on 2026-08-08: 55 sealed days held, 2026-06-09 to 2026-08-07, on a panel of 100000 domains — and across the 7 sealed days below, 21254 of those sites edited their robots.txt, 460 of them in a way that changed what a named crawler may do.
A site's robots.txt is free to read and impossible to look back at: publishers overwrite it in place, so the day a publisher started blocking a crawler is simply gone. We have held a dated daily copy across a 100,000-domain panel since 2026-06-09. Name the domains you care about and we tell you who started blocking which AI crawler, who stopped, and on which day — with the sealed snapshot behind every row.
$175/month
billed monthly, cancel any time
Request this →Checkout for this service opens by emailed invoice link. Tell us where to send it and we reply the same working day.
Anyone can read a site’s robots.txt right now. Nobody can read the one it served last month, because the file is overwritten in place and no archive of it is published anywhere. So when a publisher starts blocking your crawler, the evidence that it used not to is gone by the time you notice the traffic drop.
We fetch the same panel every day and keep what came back, which is the only way an AI-crawler policy question still has an answer a month later. You give us the domains you care about and the user-agent tokens you run; each month the sentinel reports the ones whose answer changed and the date it changed on. A domain we could not reach that day is reported as unreached, because a site we failed to fetch has not given us permission.
This is an AI-crawler policy record, not a rank tracker and not a legal opinion. The sentinel reports what a file says under the matching rule the standard defines, and it says nothing about what a publisher meant by writing it.
Across the 7 sealed days from 2026-08-01 to 2026-08-07, on a panel of 100000 domains, 21254 robots.txt files were edited and 460 of those edits changed what a named AI crawler is allowed to do.
| Sealed | Compared against | Span (days) | Domains compared | Files edited | Verdicts changed |
|---|---|---|---|---|---|
| 2026-08-01 | 2026-07-31 | 1 | 61027 | 3027 | 52 |
| 2026-08-02 | 2026-08-01 | 1 | 61033 | 2619 | 41 |
| 2026-08-03 | 2026-08-02 | 1 | 61162 | 2841 | 49 |
| 2026-08-04 | 2026-08-03 | 1 | 61167 | 3207 | 83 |
| 2026-08-05 | 2026-08-04 | 1 | 61154 | 3193 | 61 |
| 2026-08-06 | 2026-08-05 | 1 | 60947 | 3161 | 68 |
| 2026-08-07 | 2026-08-06 | 1 | 61057 | 3206 | 106 |
What moved, by crawler:
Showing 40 of 400 distinct verdict moves in that week.
Archive held: 55 sealed days, 2026-06-09 to 2026-08-07.
Days inside the archive span with no collection run, named rather than smoothed over: 2026-06-11, 2026-06-12, 2026-06-19, 2026-07-04, 2026-07-15. The longest gap ran 3 days.
Why no site is named here. No domain is named here and no robots.txt file is reproduced. Crawler policy is published per site, and thousands of those files carry contact addresses in their comments, so this feed exposes mechanical verdicts and aggregate movement only. A subscription reports the same movement against the domains you give us, which are yours to name.
One hundred thousand domains are fetched on a fixed daily schedule and each response is sealed as fetched. A day's diff compares a domain against its own previous SEALED reading, so a missed collection day widens the span rather than inventing a daily rate, and the span is printed on every row. A verdict is the RFC 9309 grouping rule applied to an exact user-agent token match: it reports what the file says, never what a site intends. A domain we could not reach has no observed policy and is excluded from both sides of every figure — unreachable is unknown, not permission.
Machine-readable offer details
US Tech Automations is a data and automation company. These are our own service terms, not a government fee, filing cost, or permit charge. Nothing on this page is legal, engineering, or financial advice.