What we can and can't see

Every census has a denominator. Ours: on 2026-09-14 we read the policy of 646 of 1020 tracked domains (63.3%). The other 374 are listed below with the reason — never guessed, never quietly dropped. Snapshot: 2026-09-14

63.3%
of tracked domains had a readable policy (646/1020)
374
unreadable — excluded from every published percentage
184
largest single cause: domain does not resolve

Every outcome, counted

OutcomeSitesShareWhat it means
robots.txt read59958.7%We fetched and parsed the site's robots.txt.
domain does not resolve18418%No DNS record exists for this name, so it never served a website. Tranco ranks domains by how often resolvers look them up, so its top list includes CDN, DNS and telemetry endpoints that answer queries but host no pages.
host answered nothing797.7%The name resolves and a host exists, but nothing served robots.txt — a timeout, TLS failure or dropped connection. Unlike a non-resolving domain, this one may well read fine tomorrow.
refused (401/403)565.5%The server answered, but refused to hand over robots.txt. We never retry with a disguised user agent.
no robots.txt (404)474.6%The site serves no robots.txt, which under RFC 9309 means every crawler is allowed.
HTML instead of text434.2%The server returned a web page where robots.txt should be — usually a catch-all router or a soft 404. We refuse to guess a policy from HTML.
HTTP 20230.3%The server answered with an unexpected HTTP status instead of the file.
HTTP 42920.2%The server answered with an unexpected HTTP status instead of the file.
HTTP 41820.2%The server answered with an unexpected HTTP status instead of the file.
HTTP 49820.2%The server answered with an unexpected HTTP status instead of the file.
HTTP 50010.1%The server answered with an unexpected HTTP status instead of the file.
HTTP 52210.1%The server answered with an unexpected HTTP status instead of the file.
HTTP 40910.1%The server answered with an unexpected HTTP status instead of the file.

Why nothing answered

Until 2026-09-14 this page had a single no HTTPS response bucket, which hid the most important distinction in the whole ledger. It is now split on recorded evidence, not on guesswork about domain names: 184 domains have no DNS record at all, and 79 resolve to a real host that then gave us nothing.

The first group is not censorship — it's the panel. Tranco ranks domains by how often the world's resolvers look them up, so its top list is full of names that were never websites: CDN edges (akam.net, tiktokcdn.com), connectivity and DNS endpoints (msftconnecttest.com, dns-parking.com), cloud plumbing (ax-msedge.net, azurewebsites.net). They have no robots.txt because they have no site, and a name that does not resolve today will almost certainly not resolve tomorrow. We leave them in the index — greyed out and marked, never deleted — so the panel stays reproducible from a public ranking instead of from our taste. The second group is the genuinely interesting one: those are live hosts, and a timeout today can be a readable policy next week, which is why coverage moves a little every day.

Network resultDomainsMeaning
dns184domain does not resolve (no such host)
timeout39no answer within 15s
tls31TLS/certificate failure
refused4connection refused
reset3connection dropped mid-request
other2other network error

How the panel changes

The other half of the denominator is which domains are in it at all. The panel was filled once from a Tranco ranking and would quietly go stale — the ranking drifts about 2% a week — so it is re-examined weekly, additions only, and capped at 25 per run so one bad list day cannot reshape the census unattended. Every run is appended to panel-history.json, and this section is that file, read at build time: 1 refresh so far, 20 domains added, 20 kept after falling out.

A domain that falls out of the ranking stays in the index. Dropping it would delete history: its daily snapshots are append-only facts about what that site's robots.txt said on those mornings, and /site/<domain>/ is a live URL somebody may have linked. So a departure is reported and kept — it keeps its page, its history and its place in the counts — and the panel grows rather than churns. Same principle as leaving the 184 no-DNS domains listed and marked: the panel is reproducible from a public ranking plus a written ledger, not from our taste about which sites deserve to be in it.

RefreshSource listPanelAddedFell out, keptDeferred
2026-09-07Tranco list 74V6X1,000 → 1,02020200
2026-09-07 — 20 added, 20 kept after falling out

Measured against Tranco list 74V6X. A newcomer has no site page until the next morning's crawl reaches it.

Added: actify.nl #103 · hostinger.com #327 · statsigapi.net #441 · bitdefender.net #819 · h264.io #833 · buydomains.com #841 · usps.com #856 · pendo.io #902 · usercontent.goog #910 · ryanair.com #913 · pccc.com #915 · dnse0.com #924 · epam.com #931 · tbank.ru #937 · zenecn.net #938 · bytefcdn.com #956 · opera-api2.com #957 · xerox.com #976 · immedia-semi.com #977 · dnse2.com #980

Fell out of the band, kept: presage.io #985 · kwai.com #986 · stbid.ru #987 · otto.de #988 · cdn-vk.ru #989 · b-msedge.net #994 · opera-api.com #995 · eu.com #998 · umich.edu #1,014 · digitaloceanspaces.com #1,071 · singtelchina.cn #1,087 · warnerbros.com #1,116 · yieldmo.com #1,170 · dyingbirds.com #1,178 · rocketmoney.dev #1,272 · tld-servers.ru #2,035 · alrosa.ru #2,616 · workbox.dk #4,141 · nominetdns.uk outside the list · nic.direct outside the list

The unreadable, by name

Listed for auditability: if you think one of these should be readable, fetch its robots.txt yourself and tell us. We never retry a refusal with a disguised user agent, and we never infer a policy from an HTML page.

domain does not resolve — 184 domains

1rx.io · 2mdn.net · 3gppnetwork.org · 3lift.com · 4dex.io · a2z.com · ad-tech.ru · adobedc.net · adtrafficquality.google · aiv-delivery.net · akam.net · akamaihd.net · akamaitech.net · alibabadns.com · alipaydns.com · aliyuncs.com · aliyuncsslbintl.com · allawnos.com · amazon.dev · app-analytics-services.com · avsxappcaptiveportal.com · aws.dev · awsglobalaccelerator.com · awswaf.com · ax-msedge.net · azure-devices.net · b-msedge.net · bamgrid.com · bdydns.com · beyondwickedmapping.org · bidr.io · bidswitch.net · browser-intake-datadoghq.com · bx-msedge.net · bytedns1.com · bytefcdn-oversea.com · bytefcdn.com · byteglb.com · byteoversea.net · capcutapi.com · cdn-apple.com · cdn-vk.ru · cdn20.com · cdnbuild.net · cloudsink.net · dbankcloud.ru · discord.media · dns-parking.com · dnse0.com · dnse2.com · dnsowl.com · dotaplabs.net · dual-s-msedge.net · dv.tech · dyingbirds.com · e2ro.com · easebar.com · edgcdn.net · enacdn.net · eu-1-id5-sync.com · exp-tas.com · facebook.net · fastly-edge.com · gandi-ns.fr · geobasket.ru · go-mpulse.net · googlezip.net · goskope.com · grammarly.io · gwfb.net · h264.io · herokudns.com · heytapdl.com · heytapmobile.com · ibyteimg.com · iiko.it · imcmdb.net · immedia-semi.com · impervadns.net · imrworldwide.com · ioref.io · ipv4only.arpa · ks-cdn.com · ksyuncdn.com · lgtvcommon.com · liadm.com · libp2p.direct · live-video.net · live.net · ln-msedge.net · media-amazon.com · microsoftonline.com · msftauth.net · msidentity.com · my.com · mynetname.net · myqcloud.com · name-services.com · namebrightdns.com · nel.goog · nflxso.net · nintendo.net · nmrodam.com · nominetdns.uk · nr-data.net · nstld.com · okcdn.ru · on.aws · onetag-sys.com · online-metrix.net · opera-api.com · opera-api2.com · orderbox-dns.com · ovscdns.com · ozone.ru · pangle.io · playstation.net · presage.io · pv-cdn.net · pvp.net · qlivecdn.com · rbxcdn.com · resolver.arpa · ripn.net · safebrowsing.apple · samsungacr.com · samsungapps.com · samsungcloudsolution.com · samsungiotcloud.com · samsungqbe.com · sc-cdn.net · sc-gw.com · service.gov.uk · sfx.ms · shopifysvc.com · singtelchina.cn · spo-msedge.net · spov-msedge.net · squarespacedns.com · ssl-images-amazon.com · static.microsoft · statsigapi.net · steamserver.net · steamstatic.com · t-msedge.net · t-s1-msedge.net · telephony.goog · tencent-cloud.net · tiktokcdn-eu.com · tiktokcdn-us.com · tiktokcdn.com · tiktokv.eu · tiktokv.us · tld-servers.ru · tm-azurefd.net · tplinknbu.com · tsyndicate.com · ttdns2.com · ttvnw.net · turn.com · ui-dns.com · usgovcloudapi.net · vecdnlb.com · vedcdnlb.com · vedsalb.com · vercel-dns-3.com · vidaahub.com · virginm.net · vkuser.net · volcfcdndvs.com · volcgslb.com · wac-msedge.net · wcdnga.com · webhostbox.net · whecloud.com · workbox.dk · wsdvs.com · xcal.tv · yahoodns.net · yandexcloud.net · yccdn.ru · yellowblue.io · zenecn.net · zpath.net

host answered nothing — 79 domains

360safe.com · a-mo.net · adgrx.com · adobe.net · ailawandorder.com · alicdn.com · avcdn.net · azurewebsites.net · beian.gov.cn · caixa.gov.br · cdngslb.com · cdnhwc1.com · chinamobile.com · cookiedatabase.org · costco.com · ddnss.de · dnspod.net · dotomi.com · douyincdn.com · edgecdn.ru · elasticbeanstalk.com · ezviz7.com · featureassets.org · firetvcaptiveportal.com · flashtalking.com · forms.gle · gamepass.com · gcdn.co · hichina.com · hm.com · hotels.com · hstgr.net · inner-active.mobi · jomodns.com · kaspersky-labs.com · keenetic.io · kroger.com · kunluncan.com · list-manage.com · lsrelayaccess.com · mangosip.ru · mcafee.com · mckinsey.com · minecraft.net · msftconnecttest.com · msftncsi.com · mybluehost.me · myhuaweicloud.com · nease.net · nflximg.com · nflxvideo.net · nic.direct · no-ip.com · npr.org · opentable.com · ovh.net · ozon.ru · prodregistryv2.org · queniuaa.com · rocket-cdn.com · run.app · sberbank.ru · sharepoint.com · shiabank.com · shifen.com · shopeemobile.com · spaceweb.pro · stbid.ru · supertms.com · tbcache.com · trueconf.net · twc.com · united.com · ups.com · wbx2.com · worldnic.com · wswebcdn.com · xboxlive.com · xiaomi.net

refused (401/403) — 56 domains

33across.com · accuweather.com · adidas.com · allaboutcookies.org · ancestry.com · apartments.com · arubanetworks.com · autodesk.com · bluehost.com · buydomains.com · cars.com · cmediahub.ru · coupang.com · dbankcloud.com · discogs.com · epam.com · epicgames.com · fandom.com · focus.de · force.com · genius.com · godaddy.com · hostgator.com · hostgator.com.br · iso.org · lenovo.com · life360.com · lijit.com · lowes.com · mayoclinic.org · mdpi.com · mhverifier.ru · mi.com · nexusmods.com · ngenix.net · nih.gov · noaa.gov · npmjs.com · oracle.com · oraclecloud.com · oup.com · politico.com · quizlet.com · sciencedirect.com · sephora.com · state.gov · substack.com · t-mobile.com · telecid.ru · theinformation.com · udemy.com · unesco.org · usda.gov · webex.com · weforum.org · xiaomi.com

HTML instead of text — 43 domains

amazonalexa.com · appspot.com · azure.com · cloud.microsoft · digitaloceanspaces.com · dropcatch.com · dyndns.org · g.page · github.io · googledomains.com · hicloud.com · icloud-content.com · id5-sync.com · internetwarriors.net · khanacademy.org · lencr.org · live.com · me.com · mobolize.com · mzstatic.com · nginx.com · nic.ru · office.com · onelink.me · outlook.com · playrix.com · reg.ru · rt.ru · schwab.com · scribd.com · skyhigh.cloud · skype.com · slideshare.net · teads.tv · telegram.me · telekom.net · timeweb.ru · ubi.com · webempresa.eu · windows.com · withgoogle.com · workers.dev · youku.com

HTTP 202 — 3 domains

aboutads.info · ieee.org · yieldmo.com

HTTP 429 — 2 domains

adjust.com · nasa.gov

HTTP 418 — 2 domains

stackoverflow.com · vkuserphoto.ru

HTTP 498 — 2 domains

wb.ru · wildberries.ru

HTTP 500 — 1 domain

eye4.cn

HTTP 522 — 1 domain

pages.dev

HTTP 409 — 1 domain

share-dns.com

Receipts for every llms.txt we report

We say 112 tracked sites publish an llms.txt. Each of those claims should be checkable by someone who is not us, so we archive the file we read. Bodies above the 256KB archive cap are kept as a SHA-256 prefix and a byte count instead — enough to prove the body existed, and enough to notice if it is silently swapped — because a claim with no evidence behind it is the failure mode this index has corrected three times.

ReceiptSitesWhat is stored
full body archived105The exact file, overwritten daily — git history keeps every version
hash + byte count7Body exceeded the 256KB archive cap; we store its SHA-256 prefix and size
no receipt yet0Recorded before receipts covered oversized files; the next crawl fills these in
Oversized llms.txt — 7 sites

datadoghq.com 20c8ac0c87cf83c6 (568KB) · mailchimp.com 8ef6cfb9a7475946 (513KB) · salesforce.com c65c1ee7df6a23f3 (998KB) · selectel.ru b6f7881928de2318 (500KB) · sourceforge.net 7b3221944bf1c5dc (1024KB) · unity3d.com 818f634428fc671c (909KB) · zendesk.com 047693a72525ecfe (724KB)

How many sites actually answered the llms.txt question

An llms.txt probe has three outcomes, and only two of them are answers. A served file is a yes; a 404 is a no; a refusal, a rate-limit, a 5xx or a timeout is no answer at all, and we record it as such rather than counting it as a no. On 2026-09-14 we got a definitive answer from 622 of 1020 tracked domains.

llms.txt probeDomainsShareWhat it means
published11211%A non-empty file was served and read — see the receipts above
none found51050%The site answered, and the answer was that there is no file
no answer39839%Refused, rate-limited, errored or timed out — 135 of these are sites whose robots.txt we could still reach

Until 2026-09-14 the index drew "no answer" and "none found" the same way — a dash — which published an unknown as a fact. The clearest case is a site we hold a receipt for: shein.com's 64KB llms.txt is archived in this repository, and on the mornings its probe is refused the old rendering said the file did not exist. Those cells now read ?, and the site page says we got no answer instead of asserting a negative. Nothing about the adoption figure changed: 11% is a floor, counted only from sites that gave us a real yes, and the 135 live sites in the last row are the room above it.

No answer today, but the site is reachable — 135 domains

33across.com · aboutads.info · accuweather.com · adidas.com · adjust.com · aliexpress.com · allaboutcookies.org · allrecipes.com · amazon.com.au · amazon.de · ancestry.com · anydesk.com · apartments.com · apnews.com · arubanetworks.com · att.net · autodesk.com · autotrader.com · bloomberg.com · bluehost.com · box.com · britannica.com · buydomains.com · cambridge.org · canva.com · cars.com · chatgpt.com · chegg.com · cmediahub.ru · comcast.net · coupang.com · creativecdn.com · dbankcloud.com · deviantart.com · discogs.com · economist.com · epam.com · epicgames.com · epicurious.com · espn.com · eye4.cn · fandom.com · fidelity.com · focus.de · force.com · ft.com · g.co · gamespot.com · genius.com · gitlab.com · glassdoor.com · godaddy.com · goo.gl · googleblog.com · herokuapp.com · hostgator.com · hostgator.com.br · ieee.org · imdb.com · indeed.com · intuit.com · investopedia.com · iso.org · lenovo.com · life360.com · lijit.com · lowes.com · marketwatch.com · markmonitor.com · mayoclinic.org · mdpi.com · medium.com · meraki.com · mhverifier.ru · mi.com · mobile.de · namecheap.com · nasa.gov · nexusmods.com · ngenix.net · nih.gov · noaa.gov · npmjs.com · office365.com · openai.com · oracle.com · oraclecloud.com · oup.com · patreon.com · people.com · perplexity.ai · pexels.com · pixabay.com · politico.com · quantserve.com · quizlet.com · reddit.com · researchgate.net · reuters.com · robinhood.com · sciencedirect.com · seekingalpha.com · sentinelone.net · sephora.com · seriouseats.com · shalltry.com · share-dns.com · shein.com · state.gov · substack.com · t-mobile.com · tandfonline.com · telecid.ru · theinformation.com · time.com · tripadvisor.com · trustpilot.com · udemy.com · umich.edu · unesco.org · unsplash.com · usatoday.com · usda.gov · vkuserphoto.ru · vrbo.com · vungle.com · washingtonpost.com · wayfair.com · wb.ru · weforum.org · wildberries.ru · wsj.com · xiaomi.com · yieldmo.com · zillow.com

How this affects the numbers

Headline rates — for example "31.9% block at least one AI crawler" — divide by 646, the readable count, never by 1020. Per-bot and per-category rates do the same. Change detection ignores any domain that was unreadable on either side of a diff, so a site going briefly unreachable never produces a fake policy flip. The raw counts behind this page are in latest.json: every domain carries its own fetch outcome.

Coverage moves day to day — timeouts and refusals are not permanent verdicts. This page is regenerated from the latest snapshot every morning, and the daily snapshots keep the history.