Blog
Analytics12 min readVisitTrack Team

Bot Traffic in Analytics: How to Filter It

How much of your analytics traffic is bots, how to recognize it (data-center cities, zero engagement, odd screens), and how to filter it without losing people.

Bot traffic in analytics is any visit recorded by a non-human client — scrapers, headless browsers, monitoring tools, AI agents and crawlers that execute your tracking script. Across the whole web, Imperva estimates automated traffic passed 53% of all requests in 2025, but only the bots that run JavaScript reach a typical analytics tool, so the share inside your dashboard is usually much smaller and much spikier. To filter it, score each visit on several signals (user agent, data-center network, browser automation flags, interaction) and exclude visits that cross a threshold, while never flagging someone who signs up or pays.

Key takeaways

  • Imperva's 2026 Bad Bot Report puts automated traffic above 53% of all web traffic in 2025, with bad bots alone at 40%.
  • Script-based analytics only sees bots that execute JavaScript — headless Chrome, Playwright and Puppeteer scrapers, some monitors and agents — which is still enough to distort small sites badly.
  • The classic fingerprints are traffic from data-center cities, zero interaction, impossible screen sizes, bursts of pageviews and user agents that don't match the browser's features.
  • No single signal is proof; VPN users and privacy browsers trip one rule at a time, so filtering should add up several signals before calling a visit a bot.
  • Bots hurt more than vanity numbers: they deflate conversion rates, inflate ad clicks and dilute revenue per visitor.

How much web traffic is bots?

The most cited measurement is Imperva's annual Bad Bot Report. The 2025 edition found automated traffic at 51% of all web traffic in 2024 — the first time in a decade it passed human traffic — with bad bots at 37%. The 2026 edition put automated traffic above 53% in 2025 and bad bots at 40%, driven largely by AI-assisted tooling that makes bots cheaper to build and harder to detect.

Year measuredAutomated share of web trafficBad bots shareSource
202451%37%Imperva Bad Bot Report 2025
202553%+40%Imperva Bad Bot Report 2026
Share of all requests Imperva observed across its network — not the share inside a JavaScript analytics tool.

Those numbers describe requests at the network edge: API calls, credential stuffing, scraping of raw HTML. Most of that never runs your tracking script, so it never shows up in a JavaScript-based dashboard at all. The bots that do reach your analytics are the ones that render pages in a real browser engine. They are a minority of all bots, but they are exactly the ones that look most like people.

Which bots show up in your analytics?

Bot typeRuns your script?What it looks like in analytics
Search crawlers that render (Googlebot, Bingbot)Yes, sometimesVisits from Google or Microsoft networks with a crawler user agent
Most AI crawlers (GPTBot, ClaudeBot, PerplexityBot)Usually notInvisible to script-based analytics; visible only in server logs
Headless scrapers (Playwright, Puppeteer, Selenium)YesData-center IPs, no interaction, rapid pageviews, sometimes a spoofed Chrome user agent
Uptime and performance monitorsSome doRegular visits to the same page at fixed intervals
Link-preview fetchers (Slack, Discord, X)RarelyA single hit on a shared URL, right after it was posted
Click fraud and ad botsYesPaid-campaign visits with near-zero engagement and odd geography
AI agents browsing on a user's behalfYesReal browsers driven by automation; the newest and hardest category
Your own end-to-end testsYesWebdriver-flagged browsers hitting staging-like paths on production

AI crawlers deserve a separate note because people expect to see them in analytics and don't. GPTBot, ClaudeBot and PerplexityBot mostly fetch raw HTML without running JavaScript, so a script tag never fires. If you want to know which AI systems are reading your content, you need server-side crawler tracking or log analysis; see AI crawler tracking for how VisitTrack does it, and the AI crawler robots.txt generator if you want to control which ones are allowed.

How do you identify bot traffic in your analytics?

Bot traffic tends to give itself away in the aggregate before it does in any single visit. These are the patterns worth checking when a number looks wrong.

  • Data-center geography. Sudden traffic from Ashburn (Virginia), Boardman (Oregon), Council Bluffs (Iowa), Frankfurt or Singapore often means cloud servers, not people — those are major AWS, Google Cloud and hosting regions.
  • Zero engagement. Visits with no scroll, no click, no key press and no mouse movement, especially in bulk, and a bounce rate on one segment near 100%.
  • Impossible devices. 0×0 or 800×600 screens, desktop Linux spikes, outdated browser versions in large numbers.
  • Bursts. Many pageviews within seconds from one visitor, or dozens of new visitors from one network within minutes.
  • Mismatches. A user agent that claims Safari or Firefox while the browser exposes Chrome-only features; a time zone that doesn't fit the IP's country (weak on its own — travelers and VPNs do this).
  • Unexplained referrers. Domains you've never heard of, sending traffic that never converts.
  • Ratios that drift. A visitor-to-signup rate that falls while signups stay flat almost always means the visitor count is being padded.

For a single suspicious visit, paste its user agent into a user-agent bot checker. It catches self-declared bots and HTTP clients instantly — but sophisticated scrapers send a perfectly normal Chrome user agent, which is why behavior and network signals matter more than the user agent alone.

Does Google Analytics filter bot traffic?

Partially. GA4 automatically excludes traffic from known bots and spiders, based on Google's own research and the IAB/ABC International Spiders and Bots List. You can't turn it off, and you can't see how much was excluded. What it doesn't catch is the hard category: headless browsers with a normal user agent, running from cloud IPs, which is most of what distorts small sites. GA4 users typically fall back on segments that exclude suspicious cities, screen resolutions or zero-engagement sessions, and on internal-traffic filters by IP. Our Google Analytics comparison covers the other differences.

How do you filter bot traffic without losing real visitors?

The hard part isn't catching bots, it's not catching people. A VPN user lives on a data-center IP. A privacy browser hides plugins and spoofs screen sizes. A fast reader leaves without scrolling. Any filter that bans a visit for one of those signals will throw away real customers. The approach that works is scoring: each signal adds weight, only the combination crosses a line.

  1. 1.Drop the unambiguous cases outright: self-declared bots, HTTP libraries (curl, python-requests, Go-http-client), headless browser user agents, and browsers reporting navigator.webdriver.
  2. 2.Verify search crawlers instead of trusting their user agent. A visit claiming to be Googlebot should pass a forward-confirmed reverse DNS lookup to googlebot.com; one that doesn't is a bot pretending.
  3. 3.Score soft signals: hosting-provider network, no interaction during the visit, impossible screen, missing browser languages, user-agent and feature mismatch, pageview bursts.
  4. 4.Decide on a threshold that no single soft signal reaches alone, so a VPN user or a quick reader stays human.
  5. 5.Re-evaluate after the visit ends. Interaction can only be judged once you know there wasn't any, so the final verdict should come a while after the last pageview.
  6. 6.Exempt converters. Anyone who signs up or pays is a customer, whatever the rules said; mark them human for good.
  7. 7.Keep the bots, flagged, instead of deleting them. You'll want to look at them when a number looks off, and to measure how much was filtered.

This is close to how VisitTrack works. Every visit is scored 0–100 at the moment it arrives and again about 30 minutes after it ends; hard rules (headless or automation user agents, HTTP clients, webdriver, unverified crawler claims) flag immediately, and soft rules add points — 40 for a hosting-provider network, 30 for no interaction (15 on phones), 40 for an impossible screen, 60 for a pageview burst — with a visit becoming a bot at 60. A VPN visitor who scrolls stays at 40, human. Signups and payments override everything. The full rule table is in the bot filtering docs.

Store the verdict, not the fingerprint

Bot detection uses browser signals that, if stored, would make a decent fingerprint. A privacy-respecting implementation uses them for the decision in memory and throws them away, keeping only human-or-bot, a score and one reason. If a tool keeps detailed device signals per visitor “for bot detection”, ask what else they're used for.

What does bot traffic do to your metrics?

MetricEffect of unfiltered botsExample
VisitorsInflatedA scraper run adds 2,000 visits in an hour
Bounce rate and engagement timeWorseBounce rate jumps from 48% to 61% on the affected pages
Conversion rateDeflatedSignup rate falls from 3.1% to 2.4% with no real change
Countries and devicesSkewed toward data-center regions and Linux desktopIreland or Virginia suddenly in your top three
Ad performancePaid clicks with no outcomeA campaign with high clicks and zero signups
Revenue per visitorDiluted$0.18 drops to $0.14 because the denominator grew
Usage-based billHigherEvents you pay for that no person generated
Illustrative figures; the direction of each effect is what holds everywhere.

Conversion rates are where bots cost real money, because decisions follow them. A founder who sees signup rate fall 20% might rewrite a landing page that was fine. Before reacting to any drop, check whether visitors rose without signups following — the bot signature — and look at the bounce rate by country and device. For funnels specifically, see SaaS conversion funnels that don't lie; for revenue metrics, revenue per visitor.

The billing line matters too if your analytics is priced per event. VisitTrack counts humans only in the dashboard, the API, reports and the bill, so a scraper run doesn't push you into the next tier. Whatever tool you use, check whether filtered bot traffic still counts toward your quota.

Should you block bots or just filter them from analytics?

They're different jobs. Filtering keeps your numbers honest; blocking protects your infrastructure and content. For most small sites, filtering is enough: the bots that run your analytics script are an annoyance, not an attack. Block at the edge (a WAF, Cloudflare's bot management, rate limits) when bots cause load, scrape content you sell, abuse signup forms or test stolen credentials — Imperva's 2026 report notes APIs and identity systems are the main targets of malicious automation.

For crawlers you may actually want — search engines, AI assistants that could cite you — don't block them by accident. Check your rules with a robots.txt tester before shipping, because a single over-broad Disallow line can remove you from AI answers as well as from Google.

What percentage of web traffic is bots?

Imperva's 2026 Bad Bot Report estimates that automated traffic accounted for more than 53% of all web traffic in 2025, with malicious bots at about 40%. The share inside a JavaScript analytics tool is much lower, because most bots never execute the tracking script.

How do I know if my website traffic is from bots?

Look for traffic from data-center cities such as Ashburn or Boardman, visits with zero scrolling or clicking, bursts of pageviews within seconds, impossible screen sizes and conversion rates that fall while conversions stay flat. Several of these together are a strong sign; any one alone is not.

Does Google Analytics 4 exclude bot traffic?

GA4 automatically excludes known bots and spiders using Google's research and the IAB/ABC International Spiders and Bots List. It doesn't reliably catch headless browsers with normal user agents, and it doesn't show how much traffic was excluded.

Do AI crawlers like GPTBot show up in analytics?

Usually not. Most AI crawlers fetch raw HTML without running JavaScript, so script-based analytics never records them. You need server-side crawler tracking or server logs to see GPTBot, ClaudeBot or PerplexityBot visits.

Can bot filtering remove real visitors?

It can if it relies on single signals. VPN users come from data-center IPs and privacy browsers hide device details. Good filtering scores several signals together, keeps a single soft signal below the bot threshold, and always treats visitors who sign up or pay as human.

Why did my traffic spike overnight with no signups?

A spike in visitors with no matching rise in signups, engagement or revenue is the most common signature of a scraper or bot run. Check the country, device and network breakdown for the spike, and compare the human-only and all-traffic views if your analytics separates them.