Glossary · SEO & AI search
What is robots.txt?
robots.txt is a plain-text file at the root of a website (/robots.txt) that tells crawlers which URLs they are allowed or not allowed to request, using the Robots Exclusion Protocol, standardized as RFC 9309 in 2022.
Also called: Robots exclusion protocol, REP
Updated
robots.txt syntax
User-agent: * Disallow: /app/ Allow: /app/public/ User-agent: GPTBot Disallow: / Sitemap: https://example.com/sitemap.xml
Rules are grouped by User-agent. Disallow and Allow take path prefixes, with * as a wildcard and $ for end of URL. The most specific matching rule wins. A crawler follows the group that best matches its name, falling back to *. Sitemap lines are independent of groups.
robots.txt for AI crawlers
| Goal | Rule |
|---|---|
| Keep content out of OpenAI training, stay in ChatGPT search | Disallow GPTBot; allow OAI-SearchBot and ChatGPT-User |
| Opt out of Gemini training without leaving Google Search | Disallow the Google-Extended token |
| Opt out of Anthropic training | Disallow ClaudeBot |
| Opt out of Common Crawl datasets | Disallow CCBot |
Generate a file for the bots you choose with the AI crawler robots.txt generator.
What robots.txt can't do
- It doesn't stop indexing. A disallowed URL can still be indexed (without content) if other pages link to it. Use a
noindexmeta tag or header — on a page crawlers are allowed to fetch. - It isn't security. It's public, and badly behaved bots ignore it. Protect private content with authentication.
- Google ignores some directives, such as
crawl-delayandnoindexlines in robots.txt. - Size limit. Google processes the first 500 KiB.
robots.txt example mistake
A team blocks /blog/ during a redesign and forgets to remove the rule. Within weeks, Search Console reports blog URLs as "Indexed, though blocked by robots.txt" with no snippets, and organic clicks fall. A one-line change fixes it — test changes with a robots.txt tester before deploying.
robots.txt and VisitTrack
robots.txt says what crawlers are allowed to do; VisitTrack's AI crawler tracking shows what they actually do — which bots fetched which paths, verified by reverse DNS where possible. That's how you check whether a block is respected, and whether allowing answer bots like ChatGPT-User leads to referred visits from AI assistants.
Frequently asked questions
Does robots.txt block AI crawlers?
It asks them not to crawl, and the major AI companies say their crawlers respect it, but compliance is voluntary. Blocking a training crawler like GPTBot doesn't necessarily block that company's live-fetch or search crawlers, which use different user agents.
Can robots.txt remove a page from Google?
No. Blocking crawling can leave a URL indexed without its content. To remove a page, allow crawling and add a noindex meta tag or X-Robots-Tag header, or use Search Console's removal tool for a temporary removal.
Related terms
- AI crawlersAI crawlers are automated bots operated by AI companies that fetch web pages — either to collect training data for models, to index content for AI search, or live, to answer a user's question in an assistant like ChatGPT, Claude or Perplexity.
- llms.txtllms.txt is a proposed standard for a Markdown file served at a website's root (/llms.txt) that gives large language models and AI tools a concise, curated summary of the site and links to its most useful content.
- Crawl budgetCrawl budget is the number of URLs on a site that a search engine crawler can and wants to crawl in a given period — determined, in Google's terms, by crawl capacity (how much your server can handle) and crawl demand (how much Google wants to recrawl your content).
- Canonical URLA canonical URL is the preferred version of a web page that search engines should index and rank when the same or very similar content is available at several URLs, usually declared with a rel="canonical" link element.
- Generative engine optimization (GEO)Generative engine optimization (GEO) is the practice of shaping a website's content and technical setup so that AI-powered answer engines — ChatGPT search, Perplexity, Google AI Overviews, Claude, Copilot — retrieve it, cite it and mention it in their generated answers.
Tools and guides
See which channels actually bring paying customers
VisitTrack is cookie-free analytics with revenue attribution built in. One script tag, no consent banner, live in two minutes. 14 days free, no card required.