Public data index › AI Crawler Configuration Review
See what your domain tells named AI crawlers—then get the edits.
For a technical SEO or web-operations lead at a content-rich publisher who has a concrete crawler-configuration question and needs a dated, reviewable record rather than a promise about what a provider will do.
/llms.txt format describes itself as a proposal for inference-time context; it is not access control. A correct file cannot prove that a crawler fetched it, honored it, indexed the site, cited a page, used content, or excluded content from training. A provider can change its documented crawler behavior after this review.This review records what one controlled domain returned; it does not prove what any crawler fetched, honored, indexed, cited, or used.
The fixed deliverable
$249 one time
After domain-control, scope, and source-suitability confirmation, we return a private dated PDF and matching CSV within 3 business days. No subscription or checkout begins here. You receive proposed text files; your team reviews and applies any change. We do not request repository, hosting, CDN, or search-console access.
- A dated matrix of the current
robots.txtgroups and directives that match the named crawlers in the agreed scope. - Status, final URL, canonical, HTML robots meta, and X-Robots-Tag observations for up to 50 current public URLs on the same controlled domain.
- An existing
/llms.txtcheck classified as genuine plaintext and spec-shaped, soft-200 HTML, empty-200, other response, or transport UNKNOWN. - An edit-ready proposed
robots.txtand, when useful, an optionalllms.txtdraft using original descriptions and links supplied or approved by the customer. - A source appendix naming the official provider pages consulted and their UTC fetch dates, plus a SHA-256 manifest of the delivered files.
What is actually measured
The observation starts at the public network boundary. We fetch the exact hostname's current /robots.txt and /llms.txt, retain status, final URL, content type, response time, body hash, and UTC observation time, and keep transport failure separate from an HTTP response. A timeout, DNS failure, blocked vantage, or unreadable body is UNKNOWN. It never becomes absent, allowed, blocked, valid, or zero.
For each named crawler, the PDF shows the matching user-agent group and the directives printed in the observed file. Specific groups are evaluated before a wildcard summary; we do not assume that a rule written for User-agent: * describes a crawler with its own group. A row distinguishes the domain's stated rule from any actual fetch seen in customer-provided logs. No log means crawler conduct is NOT MEASURED.
The URL sweep is similarly literal. A status row says which response this review received from one named vantage. The canonical and noindex columns quote normalized directive values and point to the observed URL. They do not declare that a search engine selected the canonical, removed a URL, or changed a ranking. Redirect loops, inconsistent variants, and inaccessible pages remain qualified findings rather than inferred outcomes.
Drafting boundary
The proposed robots.txt begins from the customer's current file. We preserve unrelated groups and call out collisions instead of rebuilding policy from a generic template. The default draft never adds a search-affecting Disallow. If a customer explicitly requests one, the request must name the crawler and path, and the draft carries a warning to test its effect before publication. We never infer that an AI-branded token is separate from search, or that a provider's training, search, and user-triggered retrieval agents share one policy.
The optional llms.txt is a concise navigation draft, not a copied content dump. It uses page links and original short descriptions approved for that domain. It does not reproduce article text, license third-party material, or grant permission. Because adoption remains unsettled, the draft is presented as an optional experiment with a dated format source—not as a route to visibility, citations, or model inclusion.
Why the review checks response quality
Our sealed daily crawler-policy panel gives a useful supply-side reason to inspect the bytes instead of checking only whether a path returns HTTP 200. In the 2026-08-07 seal, all 100,000 domains in the fixed Tranco panel had an attempted /llms.txt fetch. The classification counted 8,155 genuine plaintext responses, 10,588 soft-200 HTML responses, 308 empty-200 responses, 56,622 other HTTP responses, and 24,327 transport errors. Those five categories sum to the full 100,000-domain denominator.
The same seal recorded 62,134 genuine HTTP-200 /robots.txt responses, 13,736 other HTTP responses, and 24,130 transport errors, again totaling 100,000. These are our own panel observations, not provider adoption figures and not evidence that a buyer wants this service. A transport error describes what our collector could not observe; it does not say the domain was down or had no file.
A separate access-log count on this site covered 2026-07-10 00:50:52 UTC through 2026-08-09 12:45:57 UTC. It found 8 requests to /llms.txt: 5 identified as curl, 2 as the local test client, and 1 as GPTBot. The same log contained 342 requests to /robots.txt. A fetch is not a referral, enquiry, sale, or proof that the response was honored. We disclose this small count because it prevents an experimental format from being presented as established customer demand.
Source record reviewed 2026-08-09
The sources below are linked so the customer can re-check the rules at delivery time. Provider documentation can change, so the final source appendix records a fresh fetch date. If an official page is unavailable or unclear, the affected row is UNKNOWN and the draft does not guess.
- RFC 9309, Robots Exclusion Protocol — the standard explicitly separates rules from access authorization.
- Google Search Central robots.txt introduction — current search-crawler configuration context.
- OpenAI crawler documentation — current names and stated purposes for OpenAI crawlers.
- Anthropic crawler documentation — current Anthropic user-agent and robots guidance.
- The llms.txt proposal and its public repository.
- Apache License 2.0 in the proposal repository — license for that repository, not permission to copy a customer's or publisher's page text.
These documents support a configuration review; they do not allow us to predict a provider's future conduct. The report names both the documentation and the observed bytes, then keeps those two evidence types in separate columns.
What the deliverable does not say
| Question | Report label | Reason |
|---|---|---|
| Will a named crawler obey this rule? | NOT PROVEN | A file records the domain's stated instruction, not remote conduct. |
| Will this change search ranking or AI citation? | NOT MEASURED | The scope observes configuration and page responses, not ranking or citation systems. |
| Does this stop model training or all scraping? | NO OUTCOME CLAIM | robots.txt is not access control, and an unknown actor need not follow it. |
| Is use of the content licensed or lawful? | OUT OF SCOPE | The service is not legal, copyright, privacy, or compliance advice. |
| Was an unreachable page free of directives? | UNKNOWN | A failed probe cannot establish what the page returned elsewhere or later. |
We decline requests framed as “stop AI training,” “block scrapers,” “guarantee visibility,” “get cited,” or “make the site compliant.” We also decline third-party domains, login-only pages, paywall or bot-control bypasses, private data, credentialed work, bulk article copying, legal opinions, security enforcement, and deployment work. Those outcomes require evidence or authority this fixed review does not possess.
How a request proceeds
- Name one domain and the decision. Describe which crawler names or publisher-policy question matters. Do not send page bodies or secrets.
- Prove control. Place a random DNS TXT token or a token under
/.well-known/. An email address or assertion in the form is not enough. - Confirm scope and sources. We agree on up to 50 current public URLs, the named crawlers, the current official documentation, exclusions, proposed $249 price, and three-business-day delivery before payment is available.
- Receive files, then decide. The customer gets the dated PDF, CSV, draft text files, source appendix, and manifest. The customer reviews, tests, applies, rolls back, or declines every proposed change. We do not deploy it.
A good fit is a publisher with a current internal configuration question and an owner who can review text-file changes. A poor fit is anyone seeking a universal AI policy, legal assurance, a search-performance promise, a bypass, or outsourced production access.