State of the Agent Web
A weekly read on how the web's doors are opening and closing to AI — written from our own daily census data.
#5 — Blocklists that left with the plumbing
2026-09-22
Patreon and Kick switched off Cloudflare's managed AI block on the same morning. Wiley's 33 flips came from moving to a new platform, not from opening up. Weather.com's new AI policy says training crawlers stay blocked while the same edit unblocked GPTBot and Google-Extended. The block rate didn't move: 31.9% of 646, again.
#4 — Say one thing, serve another
2026-09-14
Corriere della Sera's new Content Signals header welcomes AI training while its crawler groups block every training bot it names. It and USA Today leave one kind of page open to the AI agents they block: the sponsored ones. USA Today's robots.txt comes in two versions that alternate. github.com names AI crawlers for the first time, to slow them rather than stop them. And a 33-flip day changed no block rate.
#3 — Most blocklists are not written, they are moved
2026-09-09
A CDN toggle, a copied community list and one publisher's fleet deployment - three ways an AI blocklist changes without anybody deciding anything. Plus the counter-example: one editor, one bot, one ticket number. And a caveat about our own growing denominator.
#2 — The web is blocking crawlers that barely exist
2026-08-30
Publishers are naming AI agents with almost no public footprint - FirecrawlAgent on 22 independent operators, a dataset name on 26 - while the newest bots stay the least-blocked. Plus four real policy flips, and a correction that removed half our changelog.
#1 — The bigger the brand, the higher the wall
2026-08-25
Founding census: 31% of the top 1,000 sites block an AI crawler — but among household names it's 49%, and among news sites it's 94%.