Reference
AI crawlers: who sends them, what each one is for, and how to verify them
Every crawler this scanner checks for — 92 named agents, grouped by what they are actually for. The column that matters most is the third one: the same company usually sends three different crawlers for three different purposes, and blocking one does not do what people expect the others to do.
Answer engines 14
These decide whether a chatbot can reach you, read you and quote you. Refusing one of these is the only kind of block that costs you answers.
| User-agent | What it is | How to verify it |
|---|---|---|
Meta-ExternalFetcher | Meta AI live fetch | not published — a request wearing this name cannot be checked |
MistralAI-User | Le Chat live fetch | not published — a request wearing this name cannot be checked |
DuckAssistBot | DuckDuckGo DuckAssist | published IP ranges |
YouBot | You.com index | not published — a request wearing this name cannot be checked |
OAI-SearchBot | ChatGPT search index | openai.com/searchbot.json |
ChatGPT-User | ChatGPT live fetch | openai.com/chatgpt-user.json |
Claude-SearchBot | Claude search index | published IP ranges |
Claude-User | Claude live fetch | published IP ranges |
PerplexityBot | Perplexity index | published IP ranges |
Perplexity-User | Perplexity live fetch | published IP ranges |
Gemini-Deep-Research | Gemini research agent | not published — a request wearing this name cannot be checked |
PhindBot | Phind | not published — a request wearing this name cannot be checked |
Kagibot | Kagi | not published — a request wearing this name cannot be checked |
Copilot-User | Microsoft Copilot fetch | not published — a request wearing this name cannot be checked |
Search indexes 10
Classic crawling, and the source most AI answer surfaces still draw on.
| User-agent | What it is | How to verify it |
|---|---|---|
Googlebot | Google Search / AI Overviews | reverse DNS to googlebot.com |
bingbot | Bing / Copilot | reverse DNS to search.msn.com |
Applebot | Apple / Siri | reverse DNS to applebot.apple.com |
Amazonbot | Amazon | published IP ranges |
DuckDuckBot | DuckDuckGo | not published — a request wearing this name cannot be checked |
YandexBot | Yandex | not published — a request wearing this name cannot be checked |
Baiduspider | Baidu | not published — a request wearing this name cannot be checked |
Seznambot | Seznam | not published — a request wearing this name cannot be checked |
Neevabot | Neeva | not published — a request wearing this name cannot be checked |
PetalBot | Huawei Petal | not published — a request wearing this name cannot be checked |
Training crawlers 43
They collect pages to train models. Refusing them is a policy choice with no effect on whether you appear in answers — we report it and never score it.
| User-agent | What it is | How to verify it |
|---|---|---|
cohere-ai | Cohere | not published — a request wearing this name cannot be checked |
Diffbot | Diffbot knowledge graph | not published — a request wearing this name cannot be checked |
Timpibot | Timpi index | not published — a request wearing this name cannot be checked |
Google-CloudVertexBot | Vertex AI grounding | not published — a request wearing this name cannot be checked |
GPTBot | OpenAI training + index | openai.com/gptbot.json |
ClaudeBot | Anthropic training | published IP ranges |
Google-Extended | Gemini training | reverse DNS to googlebot.com |
Applebot-Extended | Apple training | reverse DNS to applebot.apple.com |
CCBot | Common Crawl | no published ranges |
Bytespider | ByteDance | no published ranges |
meta-externalagent | Meta AI | published IP ranges |
Magpie-crawler | Magpie AI | not published — a request wearing this name cannot be checked |
img2dataset | img2dataset image corpus | not published — a request wearing this name cannot be checked |
AwarioRssBot | Awario RSS | not published — a request wearing this name cannot be checked |
AwarioSmartBot | Awario smart | not published — a request wearing this name cannot be checked |
TurnitinBot | Turnitin | not published — a request wearing this name cannot be checked |
archive.org_bot | Internet Archive | not published — a request wearing this name cannot be checked |
ia_archiver | Internet Archive (legacy) | not published — a request wearing this name cannot be checked |
meta-webindexer | Meta web index | not published — a request wearing this name cannot be checked |
omgili | Webz.io omgili | not published — a request wearing this name cannot be checked |
cohere-training-data-crawler | Cohere training corpus | not published — a request wearing this name cannot be checked |
PanguBot | Huawei PanGu | not published — a request wearing this name cannot be checked |
Ai2Bot-Dolma | Allen Institute Dolma | not published — a request wearing this name cannot be checked |
FriendlyCrawler | FriendlyCrawler ML | not published — a request wearing this name cannot be checked |
VelenPublicWebCrawler | Velen | not published — a request wearing this name cannot be checked |
MyCentralAIScraperBot | MyCentral AI | not published — a request wearing this name cannot be checked |
DeepSeekBot | DeepSeek | not published — a request wearing this name cannot be checked |
ICC-Crawler | NICT ICC | not published — a request wearing this name cannot be checked |
GoogleOther | Google non-search fetch | not published — a request wearing this name cannot be checked |
anthropic-ai | Anthropic legacy agent | not published — a request wearing this name cannot be checked |
Claude-Web | Anthropic legacy agent | not published — a request wearing this name cannot be checked |
FacebookBot | Meta legacy | not published — a request wearing this name cannot be checked |
Omgilibot | Webz.io | not published — a request wearing this name cannot be checked |
ImagesiftBot | Imagesift | not published — a request wearing this name cannot be checked |
AI2Bot | Allen Institute | not published — a request wearing this name cannot be checked |
Scrapy | Generic scraper framework | not published — a request wearing this name cannot be checked |
SemrushBot-OCOB | Semrush AI corpus | not published — a request wearing this name cannot be checked |
Applebot-Extended-Ads | Apple ads corpus | not published — a request wearing this name cannot be checked |
TikTokSpider | TikTok | not published — a request wearing this name cannot be checked |
QuillBot | QuillBot | not published — a request wearing this name cannot be checked |
Webzio-Extended | Webz.io extended | not published — a request wearing this name cannot be checked |
ProRataInc | ProRata | not published — a request wearing this name cannot be checked |
AwarioBot | Awario | not published — a request wearing this name cannot be checked |
Regional search engines 7
Worth allowing if you serve those markets, harmless otherwise. Never scored.
| User-agent | What it is | How to verify it |
|---|---|---|
Yeti | Naver | not published — a request wearing this name cannot be checked |
YoudaoBot | Youdao | not published — a request wearing this name cannot be checked |
Exabot | Exalead | not published — a request wearing this name cannot be checked |
Sogou web spider | Sogou | not published — a request wearing this name cannot be checked |
YisouSpider | Yisou | not published — a request wearing this name cannot be checked |
360Spider | 360 Search | not published — a request wearing this name cannot be checked |
Sosospider | Soso | not published — a request wearing this name cannot be checked |
Social link previews 5
They render the card when someone shares your URL. Never scored.
| User-agent | What it is | How to verify it |
|---|---|---|
Twitterbot | X link preview | not published — a request wearing this name cannot be checked |
LinkedInBot | LinkedIn link preview | not published — a request wearing this name cannot be checked |
facebookexternalhit | Facebook link preview | not published — a request wearing this name cannot be checked |
Pinterestbot | not published — a request wearing this name cannot be checked | |
Slackbot-LinkExpanding | Slack unfurl | not published — a request wearing this name cannot be checked |
SEO and research crawlers 13
Third-party tools. Blocking them affects nobody’s answers, including yours. Never scored.
| User-agent | What it is | How to verify it |
|---|---|---|
AhrefsBot | Ahrefs link index | not published — a request wearing this name cannot be checked |
SemrushBot | Semrush crawler | not published — a request wearing this name cannot be checked |
MJ12bot | Majestic link index | not published — a request wearing this name cannot be checked |
DotBot | Moz link index | not published — a request wearing this name cannot be checked |
rogerbot | Moz site crawler | not published — a request wearing this name cannot be checked |
DataForSeoBot | DataForSEO | not published — a request wearing this name cannot be checked |
BLEXBot | WebMeUp link index | not published — a request wearing this name cannot be checked |
CloudflareBrowserRenderingCrawler | Cloudflare Browser Run /crawl | not published — a request wearing this name cannot be checked |
Cloudflare-AutoRAG | Cloudflare AutoRAG | not published — a request wearing this name cannot be checked |
Peer39_Crawler | Peer39 ad context | not published — a request wearing this name cannot be checked |
AdsBot-Google | Google Ads quality | not published — a request wearing this name cannot be checked |
AmazonAdBot | Amazon Ads | not published — a request wearing this name cannot be checked |
AdIdxBot | Microsoft Ads | not published — a request wearing this name cannot be checked |
The mistake almost everyone makes
Blocking GPTBot does not remove you from ChatGPT. GPTBot collects training data. OAI-SearchBot builds the index ChatGPT search reads, and ChatGPT-User fetches your page live when someone asks about you. Those are three separate permissions, and a single Disallow aimed at the wrong one either gives away training data you meant to keep or removes you from answers you meant to appear in.
The same split applies to Google (Google-Extended is Gemini training only — it has no effect on Search or AI Overviews) and to Apple (Applebot-Extended is training only).
Why the user-agent string is a claim, not a fact
Anyone can send a request calling itself GPTBot. Verification means matching the request against the address ranges the operator publishes — OpenAI, Anthropic and Perplexity all publish theirs; Common Crawl and ByteDance do not.
An access log records the string. It does not record whether the string was earned. That is the difference between GPTBot visited you and something called itself GPTBot, and on our own sites the same name has arrived both ways in a single day.
Being allowed is not the same as being served
A permission in robots.txt is a request, not a delivery. Your edge, your CDN or a bot-defence rule can refuse a named crawler regardless of what your file says — and it usually does so silently, because nothing in your analytics reports a crawler that never arrived.
We have measured sites returning 502 to GPTBot and ClaudeBot while serving Googlebot normally, and a robots.txt answering HTTP 200 with a verification page that no crawler could parse. Neither shows up in a rank tracker, a validator or an uptime check.
How to check your own site
The fastest version takes ten seconds and needs nothing installed:
curl -sI -A "GPTBot/1.2" https://yoursite.com/robots.txt
Look at the status and the content type. A 200 that returns text/html means something is answering in place of your file.
For the full comparison — every agent above, what each one received, and how much of your page was readable text — run a free scan or read a real report first.
Related
- Which AI crawlers actually visit a small business site — observed visits, not probes
- The file that returned 200 to every crawler and could not be read
- The dataset — what we have measured across every site scanned here