Blocklists that left with the plumbing
State of the Agent Web #5 · 2026-09-22
Fifth issue. Everything below is drawn from our daily census of AI-crawler access across 1,020 top websites, and every number is reproducible from the snapshots in the public repository. This issue covers 2026-09-15 to 2026-09-22.
The headline: blocklists that left with the plumbing
A quiet week. The changelog gained 68 entries, all of them on two mornings (35 on 2026-09-16, 33 on 2026-09-18); the other six days recorded nothing. Most of what did change was AI blocks disappearing. Only one of those removals comes with a stated reason. The rest left along with a piece of infrastructure: a CDN-managed section switched off, a website moved to a new platform. A block that arrived without a decision can leave the same way, and a flip in our changelog is not by itself evidence that anyone changed their mind about AI.
Two Cloudflare blocks switched off on the same morning
On 2026-09-16, patreon.com and kick.com both deleted the same 61 lines from the top of their robots.txt: Cloudflare's managed section, with its legal preamble, its Content-Signal: search=yes,ai-train=no,use=reference line and Disallow: / for nine crawlers. On each site eight tracked crawlers moved from blocked to restricted: GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, Bytespider, meta-externalagent and Amazonbot. “Restricted”, not “allowed”, because both sites' own * rules still close some paths to every crawler.
The two sites did not end up in the same place. Kick's file now names no AI crawler at all. Patreon's own hand-written groups, further down the file and unchanged, still block PerplexityBot and Timpibot. They also carry Patreon's own Content-Signal line, search=yes, ai-input=yes, ai-train=no, in the groups for PetalBot and TikTokSpider. So Patreon still states an AI-training preference in its own words; it just no longer backs it with a block on the crawlers Cloudflare's list named.
We cannot see why either site made the change, and the file does not say. The managed section is a switch in Cloudflare's dashboard, not something an editor writes, so its removal may reflect a decision about AI crawlers or a change of CDN settings made for other reasons. Two unrelated sites doing it on the same morning is a coincidence we can record but not explain.
A replatform that reads as 33 flips
On 2026-09-18, wiley.com produced 33 changelog entries: its default * policy and all 32 tracked crawlers moved from restricted to allowed. It would be easy to write that as “a major publisher opens up to AI.” That would be wrong twice.
First, Wiley had never blocked an AI crawler in our records. Every readable day since our census began, all 32 were “restricted” only because Wiley's * group closed shopping-cart, checkout, account and tracking-parameter URLs to everyone. The eight crawlers it named were SEO and marketing bots such as AhrefsBot and MJ12bot, and none of them was an AI crawler. Second, the new file is not a new policy. The old robots.txt was 3,222 bytes of path rules. The new one was 51 bytes: a single Sitemap: line pointing at a Webflow host. The whole file was replaced, and the path rules went with it, the way a site's robots.txt often does when it moves to a new publishing system.
On 2026-09-22 the file grew to 121 bytes, adding a second sitemap and an Allow: /. That line sits above any User-agent: line, so it belongs to no group and, under the robots.txt standard, applies to nobody. No verdict moved. The episode is the mirror of issue #3's lesson: blocks that arrive without a decision can also leave without one, and a changelog counts both kinds of change the same way.
The one removal that explained itself
weather.com also changed on 2026-09-16, and unlike the others it left a note. A new comment block, headed “AI / AEO bot policy” with a ticket number (CW-9994), says it allows AI agents that fetch a page because a user asked, so that its content can be cited in AI answers. It adds that bulk AI-training and indexing crawlers remain blocked, and gives YouBot and Google-CloudVertexBot as examples.
The edit did what the comment says, and more. YouBot and Google-CloudVertexBot gained new Disallow: / groups, and Brightbot got an explicit allow (ChatGPT-User had one already). But the edit also deleted the Disallow: / groups for GPTBot, ClaudeBot, anthropic-ai, Claude-Web, Google-Extended, Applebot-Extended, PerplexityBot, meta-externalagent and cohere-ai. All nine moved from blocked to restricted, which puts them under the same * rules as any other crawler. Several of those nine are training crawlers by their operators' own descriptions: GPTBot, Google-Extended and Applebot-Extended in particular. So the comment's “remain blocked” describes the crawlers still listed below it, not the ones the edit removed.
Eight tracked crawlers are still blocked there: CCBot, Bytespider, Amazonbot, Diffbot, omgili, omgilibot, YouBot and ImagesiftBot. The same edit removed 100 Disallow: /g00/ to /g99/ lines from the * group. The file's header gives its last update as 2026-09-08. We read it successfully every day from 09-08 to 09-15 and first saw the new version on 09-16, so the date in the header is not the day it went live.
One list added, naming a dataset
Going the other way, onet.pl added fourteen Disallow: / groups to what had been a two-line file on 2026-09-16. Eight of the names are crawlers we track, and each moved from restricted to blocked: Bytespider, meta-externalagent, Amazonbot, cohere-ai, Diffbot, omgilibot, YouBot and Timpibot. The other six include Ai2Bot-Dolma, which is named after the training dataset its crawler collects for rather than after the crawler. Issue #2 found that name in the files of 26 independent operators, as an example of how these lists get written: copied from list to list. Onet's list names none of GPTBot, ClaudeBot or Google-Extended.
The cohorts, one more data point
We re-counted the three copied-blocklist cohorts from issue #3 at every daily commit in this window:
- Cloudflare Managed Content: 5 → 3 on 2026-09-16, when Patreon and Kick left. The three that remain are gamespot.com, nexusmods.com and snopes.com. One caveat: GameSpot has refused our robots.txt request (HTTP 401/403) every day since 2026-09-16. The cohorts are counted from the last file we archived for each site, so it is counted on a file we can no longer confirm.
- The robotstxt.com/ai community list: 1 site, launchpad.net, for a third straight week.
- Content Signals: 25 → 24. Kick left with Cloudflare's section. Patreon stays, on the strength of its own lines. There were no additions.
Where the numbers stand
31.9% of the 646 sites we could read on 2026-09-22 block at least one tracked AI crawler, the same as in issue #4, and 206 sites both times. The unchanged figure hides movement underneath. Kick left the blocking set by unblocking, and GameSpot left because it became unreadable. onet.pl joined, and so did stbid.ru, which blocks all 32 crawlers but only answers our request on some days. Weather.com and Patreon each still block at least one crawler, so they never left.
The per-bot order holds. CCBot and Bytespider are the most-blocked at 24.9%, then ClaudeBot (22%) and GPTBot (21.2%), all slightly lower than issue #4. The newest names are still the least blocked: Gemini-Deep-Research (11.3%) and GrokBot (10.4%).
111 of the 1,020 tracked sites (10.9%) publish an llms.txt, one fewer than issue #4. azure.com, whose 47,037-byte file we archived on 2026-08-25, has answered “no file” every day since 2026-09-16. That is a definite answer, not a failed request. Among the 626 domains that gave our probe a definite answer on 09-22, the rate is 17.7%.
How to read this issue
Cohort counts come from matching the archived robots.txt of every panel domain against the marker each cohort leaves, recomputed at each day's commit rather than remembered. They are published daily on the stats page. Flip counts come from the changelog, which is derived from append-only daily snapshots. Every file quoted above was read from our archive at the commit for the day in question.
Every named site's full record is on its own page, and any two can be put side by side on the compare page. Try Patreon against Kick.
Every number in this issue is reproducible from the committed daily snapshots. Cite "Canicrawl" with a link — data CC BY 4.0.