Key takeaways
- A user-agent is a claim, not proof. Real crawlers can be confirmed by IP range or reverse DNS; anything can claim to be Chrome or Googlebot.
- Most AI crawlers (GPTBot, ClaudeBot, PerplexityBot) don't run JavaScript, so script-based analytics never sees them — you need server logs or server-side tracking.
- Headless browsers and automation tools do run JavaScript and are the bots most likely to pollute analytics.
- The AI bot's purpose matters for robots.txt: blocking a training crawler doesn't remove you from AI answers; blocking a search crawler does.
What the checker recognizes
| Category | Examples | Runs JavaScript? | Shows up in |
|---|---|---|---|
| Human browser | Chrome, Safari, Firefox, Edge, Samsung Internet, in-app browsers (Instagram, TikTok, LinkedIn) | Yes | Analytics and logs |
| AI crawler | GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, Claude-User, PerplexityBot, Perplexity-User, Meta-ExternalAgent, Amazonbot, CCBot, Bytespider | Generally no | Server logs, CDN logs, server-side tracking |
| Search engine crawler | Googlebot, Bingbot, Applebot, DuckDuckBot, YandexBot, Baiduspider | Googlebot and Bingbot render pages, so sometimes | Logs; occasionally analytics |
| SEO tool | AhrefsBot, SemrushBot, MJ12bot, DotBot, Screaming Frog, Lighthouse | Mostly no (Lighthouse yes) | Logs |
| Link-preview bot | facebookexternalhit, Twitterbot, LinkedInBot, Slackbot, Discordbot, WhatsApp | No | Logs, the moment a link is shared |
| Monitoring | UptimeRobot, Pingdom, StatusCake, Datadog Synthetics, Checkly | Some (browser checks) | Logs; analytics for browser-based checks |
| Headless / automation | HeadlessChrome, Puppeteer, Playwright, Selenium | Yes | Analytics — the main source of fake visitors |
| HTTP library | curl, wget, python-requests, Go-http-client, node-fetch, axios | No | Logs |
How to check a user-agent
- 1.Copy the User-Agent value from your server, CDN or WAF logs (it's the quoted string near the end of a standard access-log line), or click "Use my browser's user-agent".
- 2.Paste it in. The verdict shows the category, the bot's name and operator, and for AI crawlers whether it's a training, search or user-fetch bot.
- 3.Read the VisitTrack UA rule field: that's the reason code VisitTrack's ingest would record for this string (for example ua:bot, ua:headless or ua:http-client), or none for a browser-looking string.
- 4.If it claims to be a major crawler, verify it: run a reverse-DNS lookup on the IP (Googlebot resolves to googlebot.com or google.com, Bingbot to search.msn.com) and forward-resolve the hostname back to the same IP, or match the IP against the vendor's published ranges.
Example user-agent strings
# OpenAI's training crawler
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
# Perplexity's search indexer
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)
# Googlebot smartphone
Mozilla/5.0 (Linux; Android 6.0.1; Nexus 5X Build/MMB29P) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/140.0.0.0 Mobile Safari/537.36 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)
# A headless browser (runs JavaScript, looks almost like Chrome)
Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) HeadlessChrome/140.0.0.0 Safari/537.36Why the user-agent alone isn't enough
Scrapers routinely send a current Chrome user-agent, and some "Googlebot" traffic is spoofed. That's why VisitTrack's bot filtering treats the user-agent as one signal among several. Hard rules (headless or automation user-agents, navigator.webdriver, an HTTP-library user-agent) flag a visit immediately; softer signals add up to a score — a hosting-provider IP (40 points), more than five pageviews in three seconds (60), no mouse, scroll, key or touch interaction (30), a browser that claims Safari but exposes Chrome-only features (30). A visit is a bot at 60 points. So a person on a VPN who scrolls stays human, while a datacenter visitor that never interacts gets filtered.
Common mistakes when filtering bots
- Filtering by "bot" in the user-agent. It misses headless Chrome, curl and scrapers posing as browsers. The long tail needs a maintained pattern list like isbot, plus behavioral signals.
- Trusting a Googlebot user-agent. Verify by reverse DNS before giving it special treatment (such as bypassing a paywall or rate limit).
- Blocking link-preview bots. facebookexternalhit, LinkedInBot and Slackbot need to fetch your page to build share cards; blocking them breaks every preview.
- Assuming AI traffic is in your analytics. AI crawlers don't execute your tracking script. If you want to know how often ChatGPT or Claude fetch your pages, you need server-side AI crawler tracking.
- Counting in-app browsers as bots. Instagram, TikTok and LinkedIn in-app browsers have unusual user-agents but are real people.
To decide what each AI crawler may access, use the AI crawler robots.txt generator, and test the result with the robots.txt tester.
Frequently asked questions
How can I tell if a user-agent is a bot?
Check it against a maintained list of crawler and automation patterns, which is what this tool does: self-declared bots contain tokens like Googlebot, GPTBot or AhrefsBot, and scripts identify as curl or python-requests. For bots disguised as browsers, the user-agent isn't enough; you need behavioral and network signals such as datacenter IPs and missing interaction.
What is the GPTBot user-agent?
GPTBot is OpenAI's training crawler, and its user-agent contains "GPTBot/1.x" with a link to openai.com/gptbot. OpenAI also runs OAI-SearchBot for ChatGPT search and ChatGPT-User for pages fetched on a user's request, each with its own token.
How do I verify that a request really comes from Googlebot?
Do a reverse DNS lookup on the IP address; a genuine Googlebot resolves to a hostname ending in googlebot.com or google.com. Then do a forward DNS lookup on that hostname and confirm it returns the same IP. Google also publishes its crawler IP ranges as JSON.
Do AI crawlers show up in Google Analytics?
Generally no. GPTBot, ClaudeBot, PerplexityBot and similar crawlers fetch the raw HTML without running JavaScript, so tag-based analytics like GA4 never fires for them. You can see them in server or CDN logs, or with server-side tracking such as VisitTrack's AI crawler tracking.
What does HeadlessChrome in a user-agent mean?
It means the page was loaded by Chrome running without a visible window, usually controlled by a script through Puppeteer, Playwright or Selenium. It's used for testing, scraping and some AI agents, and because it runs JavaScript it can inflate analytics unless filtered.
Is facebookexternalhit a bot I should block?
No, not if you want link previews. facebookexternalhit is the crawler Facebook, Messenger and Instagram use to read your Open Graph tags when someone shares a link. Blocking it leaves shared links without a title or image.
Related tools and guides
- AI crawler robots.txt generatorAllow or block each AI bot by purpose.
- robots.txt testerCheck which rule applies to a crawler.
- Referrer channel checkerClassify where human visits come from.
- Bot filtering in VisitTrackEvery rule and score, documented.
- AI crawler trackingLog GPTBot, ClaudeBot and Perplexity visits server-side.
- GlossaryBot, crawler and analytics terms.