CrawlCheck

Reference

AI crawlers: who sends them, what each one is for, and how to verify them

Every crawler this scanner checks for — 92 named agents, grouped by what they are actually for. The column that matters most is the third one: the same company usually sends three different crawlers for three different purposes, and blocking one does not do what people expect the others to do.

Answer engines 14

These decide whether a chatbot can reach you, read you and quote you. Refusing one of these is the only kind of block that costs you answers.

User-agentWhat it isHow to verify it
Meta-ExternalFetcherMeta AI live fetchnot published — a request wearing this name cannot be checked
MistralAI-UserLe Chat live fetchnot published — a request wearing this name cannot be checked
DuckAssistBotDuckDuckGo DuckAssistpublished IP ranges
YouBotYou.com indexnot published — a request wearing this name cannot be checked
OAI-SearchBotChatGPT search indexopenai.com/searchbot.json
ChatGPT-UserChatGPT live fetchopenai.com/chatgpt-user.json
Claude-SearchBotClaude search indexpublished IP ranges
Claude-UserClaude live fetchpublished IP ranges
PerplexityBotPerplexity indexpublished IP ranges
Perplexity-UserPerplexity live fetchpublished IP ranges
Gemini-Deep-ResearchGemini research agentnot published — a request wearing this name cannot be checked
PhindBotPhindnot published — a request wearing this name cannot be checked
KagibotKaginot published — a request wearing this name cannot be checked
Copilot-UserMicrosoft Copilot fetchnot published — a request wearing this name cannot be checked

Search indexes 10

Classic crawling, and the source most AI answer surfaces still draw on.

User-agentWhat it isHow to verify it
GooglebotGoogle Search / AI Overviewsreverse DNS to googlebot.com
bingbotBing / Copilotreverse DNS to search.msn.com
ApplebotApple / Sirireverse DNS to applebot.apple.com
AmazonbotAmazonpublished IP ranges
DuckDuckBotDuckDuckGonot published — a request wearing this name cannot be checked
YandexBotYandexnot published — a request wearing this name cannot be checked
BaiduspiderBaidunot published — a request wearing this name cannot be checked
SeznambotSeznamnot published — a request wearing this name cannot be checked
NeevabotNeevanot published — a request wearing this name cannot be checked
PetalBotHuawei Petalnot published — a request wearing this name cannot be checked

Training crawlers 43

They collect pages to train models. Refusing them is a policy choice with no effect on whether you appear in answers — we report it and never score it.

User-agentWhat it isHow to verify it
cohere-aiCoherenot published — a request wearing this name cannot be checked
DiffbotDiffbot knowledge graphnot published — a request wearing this name cannot be checked
TimpibotTimpi indexnot published — a request wearing this name cannot be checked
Google-CloudVertexBotVertex AI groundingnot published — a request wearing this name cannot be checked
GPTBotOpenAI training + indexopenai.com/gptbot.json
ClaudeBotAnthropic trainingpublished IP ranges
Google-ExtendedGemini trainingreverse DNS to googlebot.com
Applebot-ExtendedApple trainingreverse DNS to applebot.apple.com
CCBotCommon Crawlno published ranges
BytespiderByteDanceno published ranges
meta-externalagentMeta AIpublished IP ranges
Magpie-crawlerMagpie AInot published — a request wearing this name cannot be checked
img2datasetimg2dataset image corpusnot published — a request wearing this name cannot be checked
AwarioRssBotAwario RSSnot published — a request wearing this name cannot be checked
AwarioSmartBotAwario smartnot published — a request wearing this name cannot be checked
TurnitinBotTurnitinnot published — a request wearing this name cannot be checked
archive.org_botInternet Archivenot published — a request wearing this name cannot be checked
ia_archiverInternet Archive (legacy)not published — a request wearing this name cannot be checked
meta-webindexerMeta web indexnot published — a request wearing this name cannot be checked
omgiliWebz.io omgilinot published — a request wearing this name cannot be checked
cohere-training-data-crawlerCohere training corpusnot published — a request wearing this name cannot be checked
PanguBotHuawei PanGunot published — a request wearing this name cannot be checked
Ai2Bot-DolmaAllen Institute Dolmanot published — a request wearing this name cannot be checked
FriendlyCrawlerFriendlyCrawler MLnot published — a request wearing this name cannot be checked
VelenPublicWebCrawlerVelennot published — a request wearing this name cannot be checked
MyCentralAIScraperBotMyCentral AInot published — a request wearing this name cannot be checked
DeepSeekBotDeepSeeknot published — a request wearing this name cannot be checked
ICC-CrawlerNICT ICCnot published — a request wearing this name cannot be checked
GoogleOtherGoogle non-search fetchnot published — a request wearing this name cannot be checked
anthropic-aiAnthropic legacy agentnot published — a request wearing this name cannot be checked
Claude-WebAnthropic legacy agentnot published — a request wearing this name cannot be checked
FacebookBotMeta legacynot published — a request wearing this name cannot be checked
OmgilibotWebz.ionot published — a request wearing this name cannot be checked
ImagesiftBotImagesiftnot published — a request wearing this name cannot be checked
AI2BotAllen Institutenot published — a request wearing this name cannot be checked
ScrapyGeneric scraper frameworknot published — a request wearing this name cannot be checked
SemrushBot-OCOBSemrush AI corpusnot published — a request wearing this name cannot be checked
Applebot-Extended-AdsApple ads corpusnot published — a request wearing this name cannot be checked
TikTokSpiderTikToknot published — a request wearing this name cannot be checked
QuillBotQuillBotnot published — a request wearing this name cannot be checked
Webzio-ExtendedWebz.io extendednot published — a request wearing this name cannot be checked
ProRataIncProRatanot published — a request wearing this name cannot be checked
AwarioBotAwarionot published — a request wearing this name cannot be checked

Regional search engines 7

Worth allowing if you serve those markets, harmless otherwise. Never scored.

User-agentWhat it isHow to verify it
YetiNavernot published — a request wearing this name cannot be checked
YoudaoBotYoudaonot published — a request wearing this name cannot be checked
ExabotExaleadnot published — a request wearing this name cannot be checked
Sogou web spiderSogounot published — a request wearing this name cannot be checked
YisouSpiderYisounot published — a request wearing this name cannot be checked
360Spider360 Searchnot published — a request wearing this name cannot be checked
SosospiderSosonot published — a request wearing this name cannot be checked

Social link previews 5

They render the card when someone shares your URL. Never scored.

User-agentWhat it isHow to verify it
TwitterbotX link previewnot published — a request wearing this name cannot be checked
LinkedInBotLinkedIn link previewnot published — a request wearing this name cannot be checked
facebookexternalhitFacebook link previewnot published — a request wearing this name cannot be checked
PinterestbotPinterestnot published — a request wearing this name cannot be checked
Slackbot-LinkExpandingSlack unfurlnot published — a request wearing this name cannot be checked

SEO and research crawlers 13

Third-party tools. Blocking them affects nobody’s answers, including yours. Never scored.

User-agentWhat it isHow to verify it
AhrefsBotAhrefs link indexnot published — a request wearing this name cannot be checked
SemrushBotSemrush crawlernot published — a request wearing this name cannot be checked
MJ12botMajestic link indexnot published — a request wearing this name cannot be checked
DotBotMoz link indexnot published — a request wearing this name cannot be checked
rogerbotMoz site crawlernot published — a request wearing this name cannot be checked
DataForSeoBotDataForSEOnot published — a request wearing this name cannot be checked
BLEXBotWebMeUp link indexnot published — a request wearing this name cannot be checked
CloudflareBrowserRenderingCrawlerCloudflare Browser Run /crawlnot published — a request wearing this name cannot be checked
Cloudflare-AutoRAGCloudflare AutoRAGnot published — a request wearing this name cannot be checked
Peer39_CrawlerPeer39 ad contextnot published — a request wearing this name cannot be checked
AdsBot-GoogleGoogle Ads qualitynot published — a request wearing this name cannot be checked
AmazonAdBotAmazon Adsnot published — a request wearing this name cannot be checked
AdIdxBotMicrosoft Adsnot published — a request wearing this name cannot be checked

The mistake almost everyone makes

Blocking GPTBot does not remove you from ChatGPT. GPTBot collects training data. OAI-SearchBot builds the index ChatGPT search reads, and ChatGPT-User fetches your page live when someone asks about you. Those are three separate permissions, and a single Disallow aimed at the wrong one either gives away training data you meant to keep or removes you from answers you meant to appear in.

The same split applies to Google (Google-Extended is Gemini training only — it has no effect on Search or AI Overviews) and to Apple (Applebot-Extended is training only).

Why the user-agent string is a claim, not a fact

Anyone can send a request calling itself GPTBot. Verification means matching the request against the address ranges the operator publishes — OpenAI, Anthropic and Perplexity all publish theirs; Common Crawl and ByteDance do not.

An access log records the string. It does not record whether the string was earned. That is the difference between GPTBot visited you and something called itself GPTBot, and on our own sites the same name has arrived both ways in a single day.

Being allowed is not the same as being served

A permission in robots.txt is a request, not a delivery. Your edge, your CDN or a bot-defence rule can refuse a named crawler regardless of what your file says — and it usually does so silently, because nothing in your analytics reports a crawler that never arrived.

We have measured sites returning 502 to GPTBot and ClaudeBot while serving Googlebot normally, and a robots.txt answering HTTP 200 with a verification page that no crawler could parse. Neither shows up in a rank tracker, a validator or an uptime check.

How to check your own site

The fastest version takes ten seconds and needs nothing installed:

curl -sI -A "GPTBot/1.2" https://yoursite.com/robots.txt

Look at the status and the content type. A 200 that returns text/html means something is answering in place of your file.

For the full comparison — every agent above, what each one received, and how much of your page was readable text — run a free scan or read a real report first.

Related