Glossary · Tracking technology
What are AI crawlers (GPTBot, ClaudeBot, PerplexityBot)?
AI crawlers are automated bots operated by AI companies that fetch web pages — either to collect training data for models, to index content for AI search, or live, to answer a user's question in an assistant like ChatGPT, Claude or Perplexity.
Also called: AI bots, LLM crawlers, GPTBot, ClaudeBot, PerplexityBot
Updated
The main AI crawlers
| User agent | Operator | Purpose |
|---|---|---|
GPTBot | OpenAI | Model training |
OAI-SearchBot | OpenAI | Search index for ChatGPT search |
ChatGPT-User | OpenAI | Live fetch when a user asks |
ClaudeBot | Anthropic | Model training |
Claude-SearchBot / Claude-User | Anthropic | Search index / live fetch |
PerplexityBot / Perplexity-User | Perplexity | Index / live fetch |
Google-Extended | A robots.txt token controlling Gemini training use, not a separate crawler | |
Applebot-Extended | Apple | Opt-out token for Apple AI training |
CCBot | Common Crawl | Open dataset widely used for training |
Bytespider, Meta-ExternalAgent, Amazonbot | ByteDance, Meta, Amazon | Training and assistants |
Why AI crawlers matter
Live-fetch bots are the new referral funnel: when ChatGPT-User or Perplexity-User reads your page, a person is waiting for an answer that may cite or link you. Training crawlers decide whether your content shapes what models know. Seeing which bots read which pages tells you where you're being used — and whether blocking them in robots.txt would cost you visibility.
Why analytics usually can't see AI crawlers
Most AI crawlers fetch the raw HTML and never run JavaScript, so script-based analytics never fires. You see them only in server logs, CDN logs, or by forwarding requests from your backend. The traffic they eventually send — people clicking a citation in chatgpt.com or perplexity.ai — does show up, as referral or AI traffic.
AI crawler example
Your docs page on cookieless tracking gets 40 ChatGPT-User fetches in a week and 12 visits from chatgpt.com. Each fetch was someone asking ChatGPT a question that your page helped answer; 12 of them clicked through. That page is earning AI visibility — keep it current and specific. Meanwhile GPTBot crawls your changelog daily, which is training, not answers.
How VisitTrack tracks AI crawlers
VisitTrack's AI crawler tracking takes a few lines in your backend middleware (Next.js, Express, Cloudflare and others) that send each request's user agent, path and optionally IP. It classifies GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Googlebot, Bingbot, Applebot, CCBot, Bytespider, Meta's crawler, Amazonbot and others into AI answers, indexing and training, and verifies claims with forward-confirmed reverse DNS where the operator supports it. Crawls never count toward your bill. Setup in AI crawler tracking; generate rules with the AI crawler robots.txt generator.
Frequently asked questions
Should I block GPTBot?
It depends on your goals. Blocking GPTBot keeps your content out of future OpenAI training data but doesn't stop live fetches by ChatGPT-User or indexing by OAI-SearchBot, which control whether ChatGPT can cite you. Many sites allow search and answer bots while deciding separately about training bots.
Do AI crawlers show up in Google Analytics?
Usually not, because most AI crawlers don't execute JavaScript. You need server logs or backend request forwarding to see them; the human visitors they later refer do appear in analytics.
Related terms
- robots.txtrobots.txt is a plain-text file at the root of a website (/robots.txt) that tells crawlers which URLs they are allowed or not allowed to request, using the Robots Exclusion Protocol, standardized as RFC 9309 in 2022.
- llms.txtllms.txt is a proposed standard for a Markdown file served at a website's root (/llms.txt) that gives large language models and AI tools a concise, curated summary of the site and links to its most useful content.
- Generative engine optimization (GEO)Generative engine optimization (GEO) is the practice of shaping a website's content and technical setup so that AI-powered answer engines — ChatGPT search, Perplexity, Google AI Overviews, Claude, Copilot — retrieve it, cite it and mention it in their generated answers.
- Bot trafficBot traffic is any website visit made by automated software rather than a person — search crawlers, AI crawlers, uptime monitors, scrapers, headless browsers and spam bots.
- User agentA user agent is the text string a browser, app or bot sends in the User-Agent HTTP header to identify its software, version and platform to the server it's requesting from.
Tools and guides
See which channels actually bring paying customers
VisitTrack is cookie-free analytics with revenue attribution built in. One script tag, no consent banner, live in two minutes. 14 days free, no card required.