AI Crawler Access Index: September 2026 Robots.txt and llms.txt Study
A pinned 1,000-domain study: 138 of 497 evaluable robots.txt responses disallow at least one selected AI token at the root; 78 llms.txt candidates pass the checks.
On 20 September 2026, we requested robots.txt and llms.txt from 1,000 domains in a pinned Tranco list. We could evaluate root-path policy in 497 responses. Of those, 138 (27.8 percent) disallowed the root for at least one of five selected AI-related product tokens. We also found 78 llms.txt candidates that passed our format and control-response checks, or 7.8 percent of the full sample.
These are observations from one measurement location and one request identity. They do not show whether a particular AI system actually crawled, indexed, trained on, or cited a site.
The sample and the missing observations
The sample is ranks 1–1,000 of Tranco list PY96J, created on 19 September 2026 from its 21 August–19 September ranking window. It is a domain ranking that includes infrastructure domains, not a count of the most-visited website pages. Requests ran from 06:23:41 to 06:27:48 UTC on 20 September.
We received 569 HTTP 200 responses to robots.txt requests. Of these, 497 were non-HTML, within the response-size limit, and contained a user-agent directive that our parser could use. The remaining 503 sample domains are unclassified for this metric: that includes failed requests, other HTTP responses, and bodies outside this definition. We do not count them as allowing or disallowing a crawler, and we do not assume the files are absent.
Every policy percentage below uses 497 as its denominator. The llms.txt candidate percentage uses the full 1,000-domain sample. Neither estimate represents every website on the internet.
Declared policy for the root path
The five-token subset contains CCBot, ClaudeBot, GPTBot, Google-Extended, and PerplexityBot. The downloadable data includes all twelve tokens checked. The subset is a reporting convention, not a claim that all five perform the same job.
- CCBot: 122 of 497 (24.5 percent).
- ClaudeBot: 106 of 497 (21.3 percent).
- GPTBot: 103 of 497 (20.7 percent).
- Google-Extended: 94 of 497 (18.9 percent).
- PerplexityBot: 84 of 497 (16.9 percent).
At least one of the five disallowed the root in 138 responses (27.8 percent). All five did so in 55 responses (11.1 percent). A rule for the root path does not tell us the policy for every other URL. For example, an Allow rule may open a subdirectory while the root remains disallowed.
The parser combines repeated exact product-token groups, falls back to the wildcard group only when no named group matches, and resolves matching root rules by specificity with Allow winning ties. It checks the URI path /, including wildcard and end-anchor rules. It does not emulate every vendor's handling of malformed files, token variants, cached policies, or network errors.
Google-Extended is a usage-control token, not the control for Google's Search crawling. Google's AI-feature guidance identifies Googlebot as the crawling control for Search. Keep training and other usage permissions separate from search-access decisions.
What counted as an llms.txt candidate
We found 78 candidates, representing 7.8 percent of the 1,000-domain sample. A candidate had to return HTTP 200, avoid HTML and truncation, contain a Markdown level-one heading and an HTTP(S) Markdown link, and differ from a fixed nonexistent control path. A control response had to be HTTP 200, 404, or 410; a failed control leaves the result unresolved. Redirects and response hashes are recorded.
There were 460 responses where this procedure did not detect a candidate, 453 unavailable or unclassifiable llms.txt responses, and 9 possible candidates whose control check was unresolved. “Not detected” does not prove absence. The format check is a heuristic; it can miss differently formatted files and cannot exclude every dynamically generated catch-all page.
The llms.txt proposal describes a publishing convention. Our check does not certify compliance with it or demonstrate that an AI provider reads the file. Google says new AI text files and special structured data are not required for its AI Overviews or AI Mode.
Reproduce the procedure and inspect the evidence
Download the September observations and the version-two script. With Python 3, run:
python3 ai-crawler-access-index-method-v2.py --list-id PY96J --size 1000 --output results.json
The output records the sample permalink, downloaded-list hash, method hash, start and finish times, requested and final URLs, HTTP status or error class, response hashes, relevant root-matching rules, and llms.txt control results. Those fields support an audit of the reported observations. Full response bodies are not republished. A later run uses the same sample but observes websites at a different time, so its counts may differ.
Correction and the August archive
The 19 August dataset and original script remain available as historical artifacts. The earlier article overstated what the published script could reproduce: it used a changing sample download, omitted the described llms.txt control probe, and handled fewer robots.txt rule cases. It also described detected files too confidently as genuine adoption.
This release pins the sample, includes the control probe, records unresolved observations, and tests the root-policy edge cases. Because the method and sample changed, the September figures must not be subtracted from August's to claim growth or decline. The September run is the baseline for this procedure.
What to do with these findings
For a site you operate, inspect the actual robots.txt rules for the crawler or usage token you intend to control. Check important page paths as well as the homepage. Then measure search impressions, visits, and useful inquiries separately. Publishing a file or opening a crawler rule is a configuration change; visibility and business outcomes require their own evidence.
Sources and related guides
- Pinned Tranco sample: PY96J · Tranco
- Robots Exclusion Protocol · RFC Editor
- Robots.txt matching and precedence · Google
- AI features and your website · Google
- The llms.txt proposal · llmstxt.org
- September observations (JSON) · Suede AI
- Version-two measurement script · Suede AI