The first open door with a name on it
State of the Agent Web #6 · 2026-10-05
Sixth issue, two weeks after the fifth. We skipped a week, and the daily census did not: every day from 2026-09-23 to 2026-10-05 is in the archive. Everything below is drawn from our daily census of AI-crawler access across 1,020 top websites, and every number is reproducible from the snapshots in the public repository.
The headline: the first site to open the door by name
The changelog gained 94 entries in the window. 89 were crawler flips, 2 were default-policy changes and 3 were new llms.txt files. 64 of the 89 flips came from just two sites rewriting their path rules, and neither named an AI crawler. Of the other 25, nine were the most deliberate loosening we have recorded so far.
Check Point writes an invitation
On 2026-10-01, checkpoint.com added a section headed “AI Agents & LLM Crawlers” to its robots.txt: thirteen named groups, each with nothing but Allow: /. Nine of the names are crawlers we track, and each moved from restricted to allowed: GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-Web, PerplexityBot, Perplexity-User, Google-Extended and CCBot. The other four are FirecrawlAgent, ExaBot, PhindBot and AndiBot. The list mixes training crawlers (GPTBot, Google-Extended, CCBot) with agents that fetch a page because a user asked, so this is not a retrieval-only welcome. It lets in the training crawlers too.
Before the edit, these crawlers fell under Check Point's * group, which closes some paths to everyone. Afterwards they each have a group of their own, and under the robots.txt standard a crawler that finds its own group ignores *. That is why the verdict is “allowed” and not merely unchanged. The same edit added a line, LLMS: https://www.checkpoint.com/llms.txt, under a comment calling it a documentation reference. No robots.txt standard defines an LLMS: field, so crawlers that follow the standard will skip it. Check Point first published its llms.txt on 2026-09-11, so the line points to a file that already existed.
Two rewrites, 64 flips, no decision about AI
Issue #5 described wiley.com replacing its robots.txt with a 51-byte file when it moved to a new platform, which turned all 32 tracked crawlers from restricted to allowed. On 2026-09-24 the path rules came back. The new file is 694 bytes, where the old one was 3,222: carts, checkout, accounts and tracking parameters closed to everyone again, with one named block for CazoodleBot. All 32 crawlers and the default policy moved back to restricted. That is 33 entries in each direction within six days, and Wiley never blocked an AI crawler in either version.
On 2026-09-25, markmonitor.com did the same thing from the other side. Its old file was a WordPress SEO plugin's stock block, with an empty Disallow: that closes nothing. The new one lists the usual WordPress paths: admin, login, search, preview and comment-feed URLs. All 32 tracked crawlers moved from allowed to restricted. None of the new lines names a crawler. The 65th flip in this group came on 2026-10-05, when telegraph.co.uk added a User-agent: Google-Extended line to the top of a group of path rules it was tidying in the same edit. That moved Google-Extended from allowed to restricted there.
Airbnb splits training from retrieval
On 2026-09-26, airbnb.com added six Disallow: / groups. Four name crawlers we track, and each moved from restricted to blocked: GPTBot, ClaudeBot, Applebot-Extended and AI2Bot. The other two are Webzio-Extended and MistralAI-Training. GPTBot had previously shared the long list of path rules Airbnb gives to search engines. Airbnb deleted that group and replaced it with the outright block. The agents that fetch pages for a live answer keep the path rules: OAI-SearchBot, ChatGPT-User and PerplexityBot. So do meta-externalagent, anthropic-ai and cohere-ai, which are also in that group. The new blocks are aimed at training crawlers, but the line is not drawn cleanly.
Smaller moves, both ways
- Tighter: name.com added a one-line block for meta-externalagent (09-25). ign.com added 86
User-agent:lines to its blocklist (10-03). Among crawlers we track, only GrokBot changed verdict, from restricted to blocked. figma.com began blocking Bytespider by name (10-01). - Looser: aol.com replaced
Disallow: /withAllow: /plus 17 path rules for ChatGPT-User and PerplexityBot (09-24). Its blocks on CCBot, Claude-Web and others stay. trustpilot.com deleted its separateDisallow: /groups for CCBot, meta-externalagent and Meta-ExternalFetcher (10-01) and added those names to its main group of path rules. That group is long, but it no longer locks these crawlers out of the whole site. - A deleted block that never worked: seekingalpha.com deleted four
Disallow: /groups on 09-27. Three were for crawlers we track (Amazonbot, Applebot-Extended, PerplexityBot), and each moved to restricted. The fourth was a duplicatePerplexity‑Usergroup written with a non-breaking hyphen (U+2011), so it never matched anything. The correctly spelled group above it still blocks Perplexity-User.
The cohorts, unchanged
The three copied-blocklist cohorts from issue #3 did not move across the window: Cloudflare Managed Content 3, the robotstxt.com/ai list 1 (launchpad.net) and Content Signals 24. The live counts are on the stats page.
Where the numbers stand
31.7% of the 646 sites we could read on 2026-10-05 block at least one tracked AI crawler (205), down from 31.9% (206) in issue #5. Bytespider is now the single most-blocked crawler at 24.8%, ahead of CCBot at 24.5%; they were tied in issue #5. ClaudeBot (21.8%) and GPTBot (21.1%) follow. The newest names are still the least blocked: Gemini-Deep-Research (11%) and GrokBot (10.2%).
116 of the 1,020 tracked sites (11.4%) publish an llms.txt, up from 111. Three new files appeared in the changelog during the window: fwmrm.net (09-24), etsy.com (09-26) and liftoff.io (09-29). As always, this figure is a floor, because many sites give our probe no definite answer.
How to read this issue
Flip counts come from the changelog, which is derived from append-only daily snapshots. Every file quoted above was read from our archive at the commit for the day in question. A flip records what a crawler that obeys robots.txt may fetch. It does not record why the file changed, and this issue shows again that most changes are housekeeping. Compare any two sites on the compare page: try Check Point against Airbnb.
Every number in this issue is reproducible from the committed daily snapshots. Cite "Canicrawl" with a link — data CC BY 4.0.