Blog
AI search12 min readVisitTrack Team

How to Track AI Crawlers: GPTBot, ClaudeBot & More

AI crawlers skip JavaScript, so your analytics never sees them. Track GPTBot, ClaudeBot and PerplexityBot from logs, a CDN or middleware, and verify them.

To track AI crawlers you have to look at requests on the server, because GPTBot, ClaudeBot, PerplexityBot and their siblings fetch raw HTML and never execute your JavaScript analytics tag. The three practical options are parsing your server or CDN logs, using your CDN's bot analytics, or adding a few lines of middleware that report each request's user agent and path to an analytics endpoint. Whichever you pick, verify the crawler's IP, because anyone can type “GPTBot” into a user-agent header.

Key takeaways

  • JavaScript analytics cannot see AI crawlers; you need server logs, CDN data or server-side middleware.
  • AI bots fall into three groups: training crawlers (GPTBot, ClaudeBot), search indexers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) and user-triggered fetchers (ChatGPT-User, Claude-User, Perplexity-User).
  • User-triggered fetches are the closest thing to a real-time signal that someone asked an AI about your page.
  • OpenAI and Perplexity publish IP ranges as JSON; check them before trusting a user agent.
  • Google-Extended and Applebot-Extended are robots.txt tokens, not crawlers you will see in your logs.

Why doesn't Google Analytics show AI crawler traffic?

Browser analytics works by running a script in the visitor's browser that sends a beacon home. AI crawlers are HTTP clients, not browsers: they request the URL, read the HTML response, and leave. The script tag is in the HTML, but nothing executes it. Googlebot and Bingbot are exceptions, because they render pages in a headless browser and do run scripts; good analytics tools filter them out as bots (here is how bot filtering works in VisitTrack). For everything else, the only record of the visit is on your server.

Which AI crawlers should you look for?

These are the user-agent tokens the major providers document, as of October 2026. Match on the token, not the full string, because version numbers change.

TokenOperatorPurposeObeys robots.txt?Published IPs
GPTBotOpenAICollects pages that may be used to train modelsYesopenai.com/gptbot.json
OAI-SearchBotOpenAIIndexes pages for ChatGPT search resultsYesopenai.com/searchbot.json
ChatGPT-UserOpenAIFetches a page when a ChatGPT user's request needs itOpenAI says rules “may not apply”openai.com/chatgpt-user.json
ClaudeBotAnthropicTraining data collectionYesDocumented by Anthropic
Claude-SearchBotAnthropicIndexing for Claude's searchYesDocumented by Anthropic
Claude-UserAnthropicFetches pages when a Claude user asksYes, per AnthropicDocumented by Anthropic
PerplexityBotPerplexityIndexes pages for Perplexity answersYesperplexity.com/perplexitybot.json
Perplexity-UserPerplexityFetches pages for a user's queryGenerally no, per Perplexityperplexity.com/perplexity-user.json
Meta-ExternalAgentMetaTraining and AI productsYesNot published as JSON
BytespiderByteDanceTrainingDisputedNot published
CCBotCommon CrawlOpen web corpus widely used for trainingYesNot published as JSON
AmazonbotAmazonAlexa and AI servicesYesVerify by reverse DNS

Sources: OpenAI's crawler documentation, Perplexity's crawler documentation and Anthropic's help center article on its crawlers. Older Anthropic tokens (anthropic-ai, Claude-Web) are deprecated, so rules that target only them no longer match current traffic.

Two tokens you will never see in logs

Google-Extended controls whether content crawled by Googlebot can be used for Gemini training and grounding; it has no user agent of its own, and Google says it does not affect inclusion in Search or AI Overviews. Applebot-Extended works the same way for Apple's models. You set them in robots.txt, but the requests still arrive as Googlebot and Applebot.

What's the difference between training, search and user-triggered bots?

Treating every AI bot as one bucket hides the useful signal. The three groups behave differently and mean different things for your business:

  • Training crawlers (GPTBot, ClaudeBot, CCBot) crawl broadly and in bursts. They tell you your content may end up in a future model, with no link back. Cloudflare Radar's crawl-to-refer data has shown training-heavy operators crawling hundreds to thousands of pages for every visitor they refer.
  • Search indexers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) keep an index that answers are built from. Blocking them is how you disappear from those products' citations.
  • User-triggered fetchers (ChatGPT-User, Claude-User, Perplexity-User) arrive because a person asked a question and the assistant decided to read your page right now. A rising count on your pricing or docs pages is the best server-side proxy you have for AI interest in your product.

A sensible default for most SaaS sites is to allow search indexers and user-triggered fetchers, and decide on training crawlers as a business question. Our AI crawler robots.txt generator builds a file for whichever policy you choose, and the robots.txt tester checks what a given bot is allowed to fetch.

How do you find AI crawlers in your server logs?

If you run Nginx, Apache or Caddy with access logs in the common “combined” format, you can get a first picture in a minute:

# Hits per AI crawler
grep -oE "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|Meta-ExternalAgent|Bytespider|CCBot|Amazonbot" \
  /var/log/nginx/access.log | sort | uniq -c | sort -rn

# Top paths fetched by user-triggered assistants
grep -E "ChatGPT-User|Claude-User|Perplexity-User" /var/log/nginx/access.log \
  | awk '{print $7}' | sort | uniq -c | sort -rn | head -20

# Did they read robots.txt, llms.txt and sitemap.xml?
grep -E "GPTBot|ClaudeBot|PerplexityBot" /var/log/nginx/access.log \
  | grep -E "/robots.txt|/llms.txt|/sitemap.xml" | awk '{print $7}' | sort | uniq -c

Logs are the most complete source, but many modern deployments do not have them. On Vercel, Netlify or most serverless platforms, request logs are short-lived or sampled. That is where the next two options come in.

Can your CDN track AI crawlers for you?

If your site sits behind Cloudflare, its AI Crawl Control dashboard (previously called AI Audit) breaks down requests by AI operator and lets you block or allow each one. Since 1 July 2025, Cloudflare blocks known AI crawlers by default on newly added domains, so check that setting before wondering why a bot never shows up. Other CDNs and WAFs (Fastly, Akamai, AWS WAF bot control) offer bot categories too, usually on higher-tier plans. The limitation is that CDN dashboards live apart from your product analytics: you cannot line up “ChatGPT-User fetched /pricing 40 times” with “signups from chatgpt.com rose this week” in one view.

How do you track AI crawlers with middleware?

The most portable approach is a few lines in your app's request middleware that forward the user agent, path and client IP to a collection endpoint, without waiting for the reply. Here is the pattern in a Next.js 16 app, where middleware is now called Proxy and lives in proxy.ts (on Next.js 15 and earlier the same code goes in middleware.ts with a function named middleware):

// proxy.ts (project root, or src/ if you use it)
import { NextResponse, type NextFetchEvent, type NextRequest } from "next/server";

export function proxy(request: NextRequest, event: NextFetchEvent) {
  // Fire and forget: the crawler never waits on this.
  event.waitUntil(
    fetch("https://visitrack.app/api/collect-bot", {
      method: "POST",
      headers: { "Content-Type": "application/json" },
      body: JSON.stringify({
        siteId: process.env.VISITRACK_SITE_ID,
        path: request.nextUrl.pathname,
        userAgent: request.headers.get("user-agent"),
        ip: request.headers.get("x-forwarded-for")?.split(",")[0]?.trim(),
      }),
    }).catch(() => {}),
  );
  return NextResponse.next();
}

export const config = {
  // Keep robots.txt, llms.txt and sitemap.xml in scope: crawlers fetch them first.
  matcher: ["/((?!_next/static|_next/image|favicon.ico).*)"],
};

The endpoint in that example is VisitTrack's, which classifies the user agent, verifies the IP and stores crawls separately from visitor events; the AI crawler tracking docs include the same pattern for Cloudflare Workers, Express and Django. The technique is not tied to one vendor: you could write the same three fields to your own database or a log drain.

How do you set up AI crawler tracking step by step?

  1. 1.Pick one layer. Use your app's middleware if you have one; use the CDN only for static sites. Installing in both counts every crawl twice.
  2. 2.Forward every request, not just the ones you think are bots. Classify on the receiving side, so a newly launched crawler shows up without a redeploy.
  3. 3.Include robots.txt, llms.txt and sitemap.xml in the matcher. They are often excluded as “static files,” and they are the first thing crawlers fetch.
  4. 4.Send the real client IP (x-forwarded-for on most platforms, cf-connecting-ip on Cloudflare), not your server's own address, so verification can work.
  5. 5.Never block the response on the report. Use waitUntil or an unawaited promise with a short timeout.
  6. 6.Test with a fake user agent: curl -A "Mozilla/5.0 (compatible; GPTBot/1.2; +https://openai.com/gptbot)" https://your-site.com/robots.txt, then confirm the hit was recorded.
  7. 7.Review weekly: crawls per bot, top paths for user-triggered fetchers, and whether search indexers reach your important pages.

How do you verify a crawler is really GPTBot?

User agents are trivially spoofed; scrapers routinely pretend to be GPTBot or Googlebot to slip past rate limits. Two verification methods cover most cases:

  • Published IP lists. OpenAI publishes JSON files of IP ranges for GPTBot, OAI-SearchBot and ChatGPT-User; Perplexity does the same for its two agents. Fetch them daily and check whether the request IP falls inside a listed CIDR range.
  • Forward-confirmed reverse DNS. For Googlebot, Bingbot, Applebot and Amazonbot, do a reverse lookup on the IP (it should end in googlebot.com, search.msn.com and so on), then a forward lookup on that hostname, which must return the same IP.

Unverified requests that claim an AI user agent are worth keeping in a separate bucket. A large unverified “GPTBot” volume is usually a scraper, and you may want to rate-limit it. You can check how a given user-agent string is classified with our user-agent bot checker.

Does blocking training crawlers hurt your AI visibility?

Not directly, according to the providers' own documentation. OpenAI separates GPTBot (training) from OAI-SearchBot (search) precisely so that sites can opt out of training while staying eligible for ChatGPT search results; Anthropic and Perplexity draw the same line between their bots. Google states that Google-Extended has no effect on inclusion in Search, including AI Overviews. So a robots.txt that disallows GPTBot and ClaudeBot but allows OAI-SearchBot, Claude-SearchBot and PerplexityBot keeps you citable.

The indirect effect is harder to measure. Models learn which brands exist and what they do partly from training data, so a site that blocks every training crawler may be mentioned less often in answers that don't involve a live search. There is no published evidence on the size of that effect. If brand awareness in AI answers matters to you more than control over your content, allowing training crawlers is the conservative choice; if the content itself is your product, blocking them is reasonable.

Whatever you decide, write it down and check it quarterly. New bots appear every few months, and a policy that lists only last year's user agents quietly allows this year's.

What should you do with AI crawler data?

What you seeWhat it likely meansWhat to do
Search indexers never reach key pagesBlocked by robots.txt, a WAF rule or a CDN defaultCheck robots.txt and CDN bot settings
High training-crawler volume, slow pagesCrawl bursts are costing you computeRate-limit or disallow training bots; keep search bots
ChatGPT-User or Claude-User on pricing and docsPeople are asking assistants about your productMake those pages answer-first and current
User-triggered fetches but no AI referralsYou are read but not cited or clickedImprove citability; see our GEO guide
“GPTBot” from unlisted IPsSpoofed user agentTreat as a scraper

Crawls are the supply side. The demand side is the humans who click through from AI answers, which shows up in your normal analytics as referrals from chatgpt.com, perplexity.ai and others. We cover that in how to track ChatGPT and AI referral traffic, and how to earn more of it in GEO for SaaS.

Can Google Analytics track GPTBot or ClaudeBot?

No. AI crawlers fetch HTML without running JavaScript, so a browser-based tag like GA4 never fires for them. You need server logs, CDN analytics or server-side middleware to see them.

What is the difference between GPTBot and ChatGPT-User?

GPTBot crawls the web for content that may be used to train OpenAI's models. ChatGPT-User fetches a specific page because a ChatGPT user's request needs it, and OpenAI says robots.txt rules may not apply to it because the action is user-initiated.

Does PerplexityBot respect robots.txt?

Yes, according to Perplexity. Its user-triggered agent, Perplexity-User, generally ignores robots.txt because it fetches pages in response to a specific user query.

How can I tell if a GPTBot request is real?

Check the request IP against the ranges OpenAI publishes at openai.com/gptbot.json. A request that claims to be GPTBot from an IP outside those ranges is almost certainly a scraper using a fake user agent.

Should I block AI crawlers?

Block training crawlers if you don't want your content in model training; that has no direct effect on AI search citations. Blocking search indexers and user-triggered fetchers removes you from those assistants' answers, which most businesses do not want.

Do AI crawlers count toward my analytics bill?

With browser-based analytics they never appear at all. With server-side crawler tracking it depends on the vendor; VisitTrack stores crawls separately and does not bill them as events.