Open dataset
What crawlers receive, measured across 306 scans
Every scan run through the checker is counted here. No domains are published — only how often each defect occurs. Updated continuously.
| Finding | Sites | Share | |
|---|---|---|---|
AI_OPTOUT_SET | 156 | 51.0% | |
NO_LLMS_TXT | 29 | 9.5% | |
NO_SITEMAP_FOUND | 18 | 5.9% | |
ROBOTS_DISALLOW_ALL | 11 | 3.6% | |
CHALLENGE_SERVED_200 | 11 | 3.6% | |
CHALLENGE_PINNED_AT_EDGE | 11 | 3.6% | |
ROBOTS_NOT_200 | 9 | 2.9% | |
SITEMAP_BLOCKED | 8 | 2.6% | |
DECLARED_SITEMAP_BROKEN | 8 | 2.6% | |
ROBOTS_BLOCKED | 6 | 2.0% | |
STALE_CACHE_SERVED | 5 | 1.6% | |
SITEMAP_IS_HTML | 4 | 1.3% | |
ROBOTS_NO_SITEMAP | 3 | 1.0% | |
ANSWER_ENGINE_REFUSED | 2 | 0.7% | |
ROBOTS_IS_HTML | 1 | 0.3% | |
PAGE_IS_MOSTLY_CODE | 1 | 0.3% |
How that compares
A share is only meaningful against a denominator. These are population counts for the same machine-layer signals, published by BuiltWith and read on 12 August 2026. They are counts of sites where the signal was detected anywhere on the web, not a sample of ours, and they are quoted here as reference points rather than reproduced as a dataset.
| Signal | Sites, web-wide | What it means |
|---|---|---|
| AppleBot disallowed | 108,452 | The largest single AI opt-out group |
| GPTBot disallowed | 107,182 | Within 1.2% of AppleBot |
| Common Crawl disallowed | 102,665 | The corpus most training sets draw on |
| ClaudeBot disallowed | 101,917 | Within 6% of AppleBot |
llms.txt published | 70,213 | The content map, still a minority signal |
| UCP profile published | 41,525 | Agentic-commerce discovery, Jan 2026 onward |
The eight largest opt-out groups sit within 12% of each other. Eight independent operators do not get blocked in near-lockstep by eight independent decisions, so most of what looks like AI policy on the web is one copy-pasted block, one plugin default or one hosting toggle — inherited rather than authored. That is the claim this dataset exists to test, one scan at a time.
Which structured facts actually exist
The same source counts detections of individual schema types. Declaring an entity is common; declaring the specific facts an assistant is asked for is an order of magnitude rarer.
| Schema type | Sites, web-wide | The question it answers |
|---|---|---|
| Organization | 354,358 | who this is |
| PostalAddress | 108,730 | where to send someone |
| ContactPoint | 91,622 | how to reach them |
| Offer | 55,182 | what it costs |
| Product | 46,455 | what is being sold |
| GeoCoordinates | 27,993 | where it is on a map |
| LocalBusiness | 26,108 | that it is a place of business |
| OpeningHoursSpecification | 22,972 | whether they are open now |
Fewer than one site in fifteen that declares an Organization declares opening hours. “Are they open?” and “how much is it?” are among the most common things anyone asks an assistant about a business, and the fields that answer them are the ones almost nobody publishes.
These are independent detection counts, not a joint distribution. Each row is the number of sites where that type was found anywhere; nothing here says which sites overlap, and no coverage rate should be inferred by dividing one row by another. Counts read from BuiltWith on 13 August 2026.
What is in the sample
Coverage matters as much as the counts: a defect rate measured entirely on one platform is a fact about that platform. The scanner deliberately widens the sample rather than scanning the same kind of site repeatedly, and the queue below is what it has not reached yet.
Crawlers we have actually seen
Everything else on this page comes from probes — we send a request wearing a crawler’s name and record what comes back. This table is the opposite: visits nobody asked for, on 4 sites we operate, over 2 days. Verified means Cloudflare matched the request to the address ranges that crawler’s operator publishes.
| Crawler | Visits | Verified | Unverified | Most-requested path |
|---|---|---|---|---|
| Googlebot | 92 | 81 | 11 | / |
| meta-externalagent | 77 | 4 | 73 | / |
| ClaudeBot | 72 | 58 | 14 | /entitymap-sitemap.xml |
| ChatGPT-User | 64 | 3 | 61 | / |
| AhrefsBot | 59 | 59 | 0 | /sitemap_index.xml |
| GPTBot | 51 | 37 | 14 | / |
| Amazonbot | 22 | 22 | 0 | /wp-json/oembed/1.0/embed |
| Bingbot | 19 | 19 | 0 | /robots.txt |
| PerplexityBot | 13 | 0 | 13 | / |
| Bytespider | 12 | 12 | 0 | /robots.txt |
| OAI-SearchBot | 9 | 8 | 1 | /robots.txt |
| SemrushBot | 3 | 3 | 0 | /robots.txt |
| Perplexity-User | 1 | 0 | 1 | /.well-known/agentic-verify.txt |
How to read this, and how not to. 494 visits across a handful of sites is a sample, not a census — it says what reached these sites, not what the web receives. And unverified is not proof of an impostor: plenty of legitimate traffic arrives from ranges nobody publishes, and our own testing appears in these counts as unverified because it is. What the column does show is that the name in a user-agent string is a claim, and it is checkable.
| Dimension | Observed |
|---|---|
| Platform | wordpress 11 |
| Rendering | server 11 |
| Size | medium 5 · small 5 |
| Business type | local-service 5 · other 2 · publisher 4 |
| Language | en 11 |
| Queued, not yet scanned | 736 domains |
How to read these numbers
Each row counts sites, not pages, and a site is counted once per finding however many
times it was scanned. A finding is recorded only when the check actually returned an answer:
lookups that failed on our side are excluded rather than counted as clean, which is why the totals
here can be smaller than the scan count. Domains are never named, in aggregate or individually,
and any operator can exclude a domain from this dataset in one line of robots.txt
— see the policy page.