← all issues

Most blocklists are not written, they are moved

State of the Agent Web #3 · 2026-09-09

Third issue. Everything below is drawn from our daily census of AI-crawler access across 1,020 top websites, and every number is reproducible from the snapshots in the public repository. This issue covers 2026-08-30 to 2026-09-09.

The headline: most blocklists are not written, they are moved

Issue #2 showed sites naming AI crawlers that barely exist. The follow-on question is who is doing the naming — and over the past ten days the answer keeps coming back the same way: a robots.txt rule usually arrives as a block of text somebody moved, not as a line somebody wrote. We now have three visible varieties of that, and one clean counter-example.

Variety one: the CDN block (a dashboard toggle)

Cloudflare can insert and maintain a managed AI section inside a customer's robots.txt, bounded by marker comments. Removing it is a toggle; nobody re-reads the user agents inside. We count the marker across the whole archived corpus every day, and it has not moved in a week: 5 sites on 2026-09-02, 5 on 2026-09-05, 5 todaygamespot.com, kick.com, nexusmods.com, patreon.com and snopes.com. The same members each day. What makes the number worth watching is what happened just before this series started: on 2026-09-01 roblox.com deleted its managed block and 8 tracked crawlers went from blocked to allowed in one move. Roblox blocks none of our 32 today.

Variety two: the copied community list

The other flavour is a roster of AI user agents copied wholesale from a community source and credited in a comment. The robotstxt.com/ai list is down to a single site in our panel — launchpad.net — and has held at one for a week. It lost its other member spectacularly: on 2026-09-02 semafor.com deleted a 96-line copied blocklist and un-blocked 31 of the 32 crawlers we track in a single commit, leaving GPTBot as its one remaining hand-written block. One deletion, thirty-one policy reversals.

Cloudflare's Content Signals preamble — the machine-readable Content-Signal: search / ai-input / ai-train header, which states a preference and blocks nothing by itself — is the one cohort with movement, and it moved down: 23 sites on 2026-09-02, 22 from 2026-09-04 onward, after stackadapt.com dropped the header. Zero sites adopted it in that window.

A denominator caveat we would rather state than hide

Our own count for that cohort reads 24 today, not 22 — and the two extra sites did not adopt anything. On 2026-09-07 the panel's weekly Tranco refresh added 20 domains, taking the census from 1,000 to 1,020, and two of the newcomers (bitdefender.net, hostinger.com) were already carrying the header when we first read them on 2026-09-08. Like for like, on the sites we were watching both days, the cohort went 23 → 22. This is the same trap as issue #2's Amazon storefronts: a count can move because the web moved, or because our lens did, and only one of those is news. Every panel change is itemised on our coverage page.

Variety three: one edit, four magazines

On 2026-09-08, newyorker.com, wired.com, epicurious.com and bonappetit.com all moved YouBot from restricted to explicitly blocked on the same morning. The four files took the identical edit — 33 lines added, 24 removed in each — which is a fleet deployment, not four editorial decisions. All four are Condé Nast titles, and all four now block 20 of the 32 crawlers we track.

The detail that makes it a story rather than a data point: arstechnica.com, also Condé Nast, has named YouBot in a Disallow: / group on every single day our probe reached it since the census began on 2026-08-25 (fifteen of sixteen days; on 2026-09-01 its robots.txt was unreadable to us). The fleet did not decide anything on Monday; it caught up with a file one of its properties was already running.

The counter-example: one human, one bot, one ticket

Against all of that, here is what an actual decision looks like. webmd.com's robots.txt carries a changelog in its header — ours records it moving from ticket CONSFE-362 to CONSFE-458 — and on 2026-09-01 exactly one thing changed: ChatGPT-User, the agent that fetches a page because a person asked a question about it, stopped being blocked. GPTBot, ClaudeBot and CCBot — the three training crawlers it names — are still blocked today, and they are the only three of our 32 it blocks. Somebody read the list, drew a line between a user's request and a training run, and filed a ticket number for it. That is rarer in this data than any cohort deletion.

Two manufacturers, opposite directions, same morning

Also on 2026-09-08: intel.com moved GPTBot, ChatGPT-User, ClaudeBot, Google-Extended, PerplexityBot and anthropic-ai from restricted to explicitly allowed, and the next morning published a 12,892-byte llms.txt — a site actively courting AI readers, blocking none of our 32. The same day, lg.com moved six of the same names the other way, from allowed to restricted. Two large hardware manufacturers, one morning, opposite conclusions about the same question.

Where the numbers stand

The aggregate has barely moved since issue #2: 31.6% of the 649 sites we could read on 2026-09-09 block at least one tracked AI crawler, against 32% of 638 on 2026-08-30. The per-bot order is unchanged too — CCBot most-blocked at 25.1%, then Bytespider (24.8%), ClaudeBot (22.5%) and GPTBot (21.7%), with the newest names still lowest: Gemini-Deep-Research (11.2%) and GrokBot (10.3%). Ten days of visible churn, and the wall is the same height. Cohorts move loudly; the average holds still.

On the other axis, 112 sites publish an llms.txt. We report that as a floor of 11% over the 624 domains that gave our probe a definitive answer, because 396 refused, timed out or errored — a non-answer is not a no, and since 2026-09-05 our pages say so explicitly rather than drawing a dash.

How to read this issue

Cohort counts come from matching the archived robots.txt of every readable domain against the marker each cohort leaves behind, re-run at build time and published on the stats page as “Blocklists nobody wrote.” Historical cohort sizes were recomputed from the committed archive at each day's commit rather than remembered. Flip counts come from the changelog, which is derived from append-only daily snapshots — we correct the derivation, never the snapshots.

Every number in this issue is reproducible from the committed daily snapshots. Cite "Canicrawl" with a link — data CC BY 4.0.