Glossary · Tracking technology

What are AI crawlers (GPTBot, ClaudeBot, PerplexityBot)?

AI crawlers are automated bots operated by AI companies that fetch web pages — either to collect training data for models, to index content for AI search, or live, to answer a user's question in an assistant like ChatGPT, Claude or Perplexity.

Also called: AI bots, LLM crawlers, GPTBot, ClaudeBot, PerplexityBot

Updated

The main AI crawlers

User agentOperatorPurpose
GPTBotOpenAIModel training
OAI-SearchBotOpenAISearch index for ChatGPT search
ChatGPT-UserOpenAILive fetch when a user asks
ClaudeBotAnthropicModel training
Claude-SearchBot / Claude-UserAnthropicSearch index / live fetch
PerplexityBot / Perplexity-UserPerplexityIndex / live fetch
Google-ExtendedGoogleA robots.txt token controlling Gemini training use, not a separate crawler
Applebot-ExtendedAppleOpt-out token for Apple AI training
CCBotCommon CrawlOpen dataset widely used for training
Bytespider, Meta-ExternalAgent, AmazonbotByteDance, Meta, AmazonTraining and assistants

Why AI crawlers matter

Live-fetch bots are the new referral funnel: when ChatGPT-User or Perplexity-User reads your page, a person is waiting for an answer that may cite or link you. Training crawlers decide whether your content shapes what models know. Seeing which bots read which pages tells you where you're being used — and whether blocking them in robots.txt would cost you visibility.

Why analytics usually can't see AI crawlers

Most AI crawlers fetch the raw HTML and never run JavaScript, so script-based analytics never fires. You see them only in server logs, CDN logs, or by forwarding requests from your backend. The traffic they eventually send — people clicking a citation in chatgpt.com or perplexity.ai — does show up, as referral or AI traffic.

AI crawler example

Your docs page on cookieless tracking gets 40 ChatGPT-User fetches in a week and 12 visits from chatgpt.com. Each fetch was someone asking ChatGPT a question that your page helped answer; 12 of them clicked through. That page is earning AI visibility — keep it current and specific. Meanwhile GPTBot crawls your changelog daily, which is training, not answers.

How VisitTrack tracks AI crawlers

VisitTrack's AI crawler tracking takes a few lines in your backend middleware (Next.js, Express, Cloudflare and others) that send each request's user agent, path and optionally IP. It classifies GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Googlebot, Bingbot, Applebot, CCBot, Bytespider, Meta's crawler, Amazonbot and others into AI answers, indexing and training, and verifies claims with forward-confirmed reverse DNS where the operator supports it. Crawls never count toward your bill. Setup in AI crawler tracking; generate rules with the AI crawler robots.txt generator.

Frequently asked questions

Should I block GPTBot?

It depends on your goals. Blocking GPTBot keeps your content out of future OpenAI training data but doesn't stop live fetches by ChatGPT-User or indexing by OAI-SearchBot, which control whether ChatGPT can cite you. Many sites allow search and answer bots while deciding separately about training bots.

Do AI crawlers show up in Google Analytics?

Usually not, because most AI crawlers don't execute JavaScript. You need server logs or backend request forwarding to see them; the human visitors they later refer do appear in analytics.

Related terms

Tools and guides

See which channels actually bring paying customers

VisitTrack is cookie-free analytics with revenue attribution built in. One script tag, no consent banner, live in two minutes. 14 days free, no card required.