The bigger the brand, the higher the wall
State of the Agent Web #1 · 2026-08-25
This is the first issue of a weekly digest drawn from our daily census of AI-crawler access across 1,000 top websites. Everything below comes from the founding snapshot of 2026-08-25; every number is reproducible from the public repository.
The headline: walls scale with fame
Across the readable top-1,000, 31.4% of sites block at least one of the 16 AI crawlers we track. Restrict the view to 232 hand-picked household names and it jumps to 48.5%. Restrict it to news sites and it hits 93.9% — the single most walled-off category on the web. The pattern is unmistakable: the more valuable a brand believes its content to be, the higher the wall it builds. The long tail of the web, meanwhile, remains largely open — most sites have simply never written an AI bot into their robots.txt.
Who's least welcome
CCBot — Common Crawl, whose open corpus historically fed many training runs — is the most-blocked bot at 25.5%, followed by ByteDance's Bytespider (25.0%), which has a reputation for ignoring the sign anyway. Notably, ClaudeBot (22.9%) is blocked more often than GPTBot (21.8%). Sites are consistently more tolerant of on-demand fetchers (a bot retrieving a page because a user asked) than of training crawlers — OAI-SearchBot is the least-blocked bot we track.
The quiet rise of the welcome mat
While the walls go up, 108 of our 1,000 sites — 10.8% — now publish an llms.txt, a file written specifically to guide AI readers. What surprised us is who: not just Cloudflare, Stripe, and Vercel, but Fox News, Target, Shein, and American Express. Welcoming AI readers has quietly stopped being a tech-company quirk and started being a mainstream commercial decision.
The default-deny genre
Eighteen sites — Reuters among them — run robots.txt files that ban every crawler they haven't explicitly invited. For these sites, any new AI bot is blocked the moment it's born, no policy update required.
Oddities
Stack Overflow answers our crawler with HTTP 418, the I'm-a-teapot joke status code. The New York Times names and bans 14 of our 16 tracked bots individually — the most thorough blocklist in the panel.
What's next
As of this week we also archive the raw robots.txt of every readable site daily (588 files in the first pass), which means from now on we can show the exact line a site changed, the day it changed. The change feed is live on the changelog and its RSS.
Every number in this issue is reproducible from the committed daily snapshots. Cite "Canicrawl" with a link — data CC BY 4.0.