Glossary · SEO & AI search
What is crawl budget?
Crawl budget is the number of URLs on a site that a search engine crawler can and wants to crawl in a given period — determined, in Google's terms, by crawl capacity (how much your server can handle) and crawl demand (how much Google wants to recrawl your content).
Also called: Crawl rate, Crawl capacity
Updated
How crawl budget is determined
Crawl budget ≈ min(crawl capacity limit, crawl demand)
Crawl capacity rises when your server responds quickly and falls with errors and slow responses. Crawl demand depends on how popular and how fresh your URLs are, and on how many URLs Google knows about. Duplicate and low-value URLs consume both.
Who needs to care about crawl budget?
Google's guidance says crawl budget mainly matters for large sites (over a million unique pages changing about weekly) and medium sites (over 10,000 pages changing daily), plus sites with many URLs reported as "Discovered – currently not indexed." A typical SaaS marketing site with a few hundred or a few thousand pages doesn't need to manage it; keeping sitemaps current is enough.
What wastes crawl budget
- Faceted navigation and filter parameters creating near-infinite URL combinations.
- Session ids and tracking parameters in internal links.
- Long redirect chains and soft 404s.
- Duplicate content without consistent canonical URLs.
- Slow server responses and 5xx errors, which lower crawl capacity.
Crawl budget example
An e-commerce site with 50,000 products generates 2 million filter URLs (?color=…&size=…). Googlebot spends most of its crawling on filter pages, and new products take weeks to be indexed. Blocking filter parameters in robots.txt and canonicalizing variants to the product page frees crawling for the pages that matter.
Checking crawl activity
Search Console's Crawl Stats report shows Googlebot's requests, response times and status codes. For a wider view, VisitTrack's AI crawler tracking records requests from Googlebot, Bingbot, Applebot and AI crawlers by path, with verified crawler identity — useful to see which sections crawlers spend their time on and whether a new section is being discovered. Crawls never count toward your VisitTrack bill.
Frequently asked questions
Does crawl budget affect small websites?
Rarely. Google says crawl budget is mainly a concern for sites with tens of thousands to millions of frequently changing URLs. Small sites are usually crawled fully without any special effort.
How do I increase my crawl budget?
Make your server fast and reliable, remove or consolidate duplicate and low-value URLs, fix redirect chains and errors, and keep sitemaps up to date. Popular, frequently updated pages also get crawled more.
Related terms
- robots.txtrobots.txt is a plain-text file at the root of a website (/robots.txt) that tells crawlers which URLs they are allowed or not allowed to request, using the Robots Exclusion Protocol, standardized as RFC 9309 in 2022.
- Canonical URLA canonical URL is the preferred version of a web page that search engines should index and rank when the same or very similar content is available at several URLs, usually declared with a rel="canonical" link element.
- Google Search ConsoleGoogle Search Console is Google's free service that shows how a website performs in Google Search — the queries it appears for, its clicks, impressions, click-through rate and average position — and reports on crawling, indexing and technical issues.
- AI crawlersAI crawlers are automated bots operated by AI companies that fetch web pages — either to collect training data for models, to index content for AI search, or live, to answer a user's question in an assistant like ChatGPT, Claude or Perplexity.
- Structured dataStructured data is machine-readable markup added to a web page — usually Schema.org vocabulary written as JSON-LD — that describes what the page contains (an article, product, organization, FAQ, defined term) so search engines and other systems can understand it precisely.
Tools and guides
See which channels actually bring paying customers
VisitTrack is cookie-free analytics with revenue attribution built in. One script tag, no consent banner, live in two minutes. 14 days free, no card required.