CrawlCheck

Open dataset

What crawlers receive, measured across 306 scans

Every scan run through the checker is counted here. No domains are published — only how often each defect occurs. Updated continuously.

FindingSitesShare 
AI_OPTOUT_SET15651.0%
NO_LLMS_TXT299.5%
NO_SITEMAP_FOUND185.9%
ROBOTS_DISALLOW_ALL113.6%
CHALLENGE_SERVED_200113.6%
CHALLENGE_PINNED_AT_EDGE113.6%
ROBOTS_NOT_20092.9%
SITEMAP_BLOCKED82.6%
DECLARED_SITEMAP_BROKEN82.6%
ROBOTS_BLOCKED62.0%
STALE_CACHE_SERVED51.6%
SITEMAP_IS_HTML41.3%
ROBOTS_NO_SITEMAP31.0%
ANSWER_ENGINE_REFUSED20.7%
ROBOTS_IS_HTML10.3%
PAGE_IS_MOSTLY_CODE10.3%

How that compares

A share is only meaningful against a denominator. These are population counts for the same machine-layer signals, published by BuiltWith and read on 12 August 2026. They are counts of sites where the signal was detected anywhere on the web, not a sample of ours, and they are quoted here as reference points rather than reproduced as a dataset.

SignalSites, web-wideWhat it means
AppleBot disallowed108,452The largest single AI opt-out group
GPTBot disallowed107,182Within 1.2% of AppleBot
Common Crawl disallowed102,665The corpus most training sets draw on
ClaudeBot disallowed101,917Within 6% of AppleBot
llms.txt published70,213The content map, still a minority signal
UCP profile published41,525Agentic-commerce discovery, Jan 2026 onward

The eight largest opt-out groups sit within 12% of each other. Eight independent operators do not get blocked in near-lockstep by eight independent decisions, so most of what looks like AI policy on the web is one copy-pasted block, one plugin default or one hosting toggle — inherited rather than authored. That is the claim this dataset exists to test, one scan at a time.

Which structured facts actually exist

The same source counts detections of individual schema types. Declaring an entity is common; declaring the specific facts an assistant is asked for is an order of magnitude rarer.

Schema typeSites, web-wideThe question it answers
Organization354,358who this is
PostalAddress108,730where to send someone
ContactPoint91,622how to reach them
Offer55,182what it costs
Product46,455what is being sold
GeoCoordinates27,993where it is on a map
LocalBusiness26,108that it is a place of business
OpeningHoursSpecification22,972whether they are open now

Fewer than one site in fifteen that declares an Organization declares opening hours. “Are they open?” and “how much is it?” are among the most common things anyone asks an assistant about a business, and the fields that answer them are the ones almost nobody publishes.

These are independent detection counts, not a joint distribution. Each row is the number of sites where that type was found anywhere; nothing here says which sites overlap, and no coverage rate should be inferred by dividing one row by another. Counts read from BuiltWith on 13 August 2026.

What is in the sample

Coverage matters as much as the counts: a defect rate measured entirely on one platform is a fact about that platform. The scanner deliberately widens the sample rather than scanning the same kind of site repeatedly, and the queue below is what it has not reached yet.

Crawlers we have actually seen

Everything else on this page comes from probes — we send a request wearing a crawler’s name and record what comes back. This table is the opposite: visits nobody asked for, on 4 sites we operate, over 2 days. Verified means Cloudflare matched the request to the address ranges that crawler’s operator publishes.

CrawlerVisitsVerifiedUnverifiedMost-requested path
Googlebot928111/
meta-externalagent77473/
ClaudeBot725814/entitymap-sitemap.xml
ChatGPT-User64361/
AhrefsBot59590/sitemap_index.xml
GPTBot513714/
Amazonbot22220/wp-json/oembed/1.0/embed
Bingbot19190/robots.txt
PerplexityBot13013/
Bytespider12120/robots.txt
OAI-SearchBot981/robots.txt
SemrushBot330/robots.txt
Perplexity-User101/.well-known/agentic-verify.txt

How to read this, and how not to. 494 visits across a handful of sites is a sample, not a census — it says what reached these sites, not what the web receives. And unverified is not proof of an impostor: plenty of legitimate traffic arrives from ranges nobody publishes, and our own testing appears in these counts as unverified because it is. What the column does show is that the name in a user-agent string is a claim, and it is checkable.

DimensionObserved
Platformwordpress 11
Renderingserver 11
Sizemedium 5 · small 5
Business typelocal-service 5 · other 2 · publisher 4
Languageen 11
Queued, not yet scanned736 domains

How to read these numbers

Each row counts sites, not pages, and a site is counted once per finding however many times it was scanned. A finding is recorded only when the check actually returned an answer: lookups that failed on our side are excluded rather than counted as clean, which is why the totals here can be smaller than the scan count. Domains are never named, in aggregate or individually, and any operator can exclude a domain from this dataset in one line of robots.txt — see the policy page.

Run a scan