Key takeaways
- AI bots come in three kinds. Blocking training crawlers (GPTBot, ClaudeBot, Google-Extended, CCBot) opts you out of future model training without removing you from AI answers.
- Blocking AI search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) removes you from those products' citations — usually not what a business wants.
- User-triggered fetchers (ChatGPT-User, Perplexity-User, Meta-ExternalFetcher) may ignore robots.txt by design, because a person asked for the page.
- Google-Extended and Applebot-Extended are robots.txt tokens, not crawlers. Blocking them doesn't affect Google Search, AI Overviews, Siri or Spotlight.
- robots.txt is a request, not enforcement. Use a WAF or bot rules for crawlers that don't comply.
Training vs search vs user-fetch bots: what blocking each one does
Most "block AI" robots.txt templates lump every AI bot together, which means site owners who only wanted to opt out of training also vanish from ChatGPT search and Perplexity citations. The vendors now separate these jobs into different user-agent tokens, so you can choose per purpose:
| Purpose | What it does | If you block it |
|---|---|---|
| Model training | Collects pages for training datasets. | Future models won't learn from new crawls of your content. Today's AI answers and citations are unaffected. |
| AI search index | Builds the index an assistant cites from (ChatGPT search, Claude search, Perplexity). | You largely drop out of that product's answers and source links. |
| User-triggered fetch | Fetches one page live because a user asked about it or pasted the link. | Vendors that honor it stop fetching; several say robots.txt may not apply to user-initiated requests. |
Every AI user-agent token in the generator (October 2026)
| Token | Provider | Purpose | Honors robots.txt? |
|---|---|---|---|
| GPTBot | OpenAI | Training | Yes |
| OAI-SearchBot | OpenAI | Search | Yes |
| ChatGPT-User | OpenAI | User fetch | Not always (user-initiated) |
| ClaudeBot | Anthropic | Training | Yes |
| Claude-SearchBot | Anthropic | Search | Yes |
| Claude-User | Anthropic | User fetch | Yes |
| anthropic-ai (legacy) | Anthropic | Training | Yes |
| PerplexityBot | Perplexity | Search | Yes |
| Perplexity-User | Perplexity | User fetch | Generally no |
| Google-Extended (token only) | Training | Yes | |
| Google-CloudVertexBot | Search | Yes | |
| Applebot-Extended (token only) | Apple | Training | Yes |
| Meta-ExternalAgent | Meta | Training | Yes |
| Meta-ExternalFetcher | Meta | User fetch | Not always (user-initiated) |
| Amazonbot | Amazon | Training | Yes |
| Amzn-SearchBot | Amazon | Search | Yes |
| Amzn-User | Amazon | User fetch | Yes |
| CCBot | Common Crawl | Training | Yes |
| Bytespider | ByteDance | Training | Reported not to |
| DuckAssistBot | DuckDuckGo | User fetch | Yes |
| MistralAI-User | Mistral | User fetch | Yes |
| cohere-ai | Cohere | User fetch | Yes |
| cohere-training-data-crawler | Cohere | Training | Yes |
| AI2Bot | Allen Institute for AI | Training | Yes |
| Diffbot | Diffbot | Training | Yes |
| omgili | Webz.io | Training | Yes |
| Timpibot | Timpi | Training | Yes |
How to use the generator
- 1.Pick a preset. "Stay in AI answers, opt out of training" is the usual choice for companies that want to be cited but not trained on.
- 2.Adjust individual bots. Unticking a bot blocks it; the badges show which ones are token-only, legacy, or may not honor the file.
- 3.Add paths every crawler should avoid (admin, API, checkout) and your sitemap URL.
- 4.Copy or download the file and publish it at https://yourdomain.com/robots.txt. If you already have one, merge the AI groups into it — a robots.txt with two User-agent: * groups still works (they're combined), but it's harder to read.
- 5.Check the result with the robots.txt tester: paste the file, a URL and a bot token, and it shows the rule that decides access.
Example: opt out of training, stay visible in AI search
# Block AI model training bots
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: Google-Extended
User-agent: Applebot-Extended
User-agent: CCBot
User-agent: Meta-ExternalAgent
Disallow: /
# Everyone else, including OAI-SearchBot, Claude-SearchBot and PerplexityBot
User-agent: *
Disallow: /admin
Sitemap: https://example.com/sitemap.xmlCommon mistakes with AI crawler rules
- Blocking Google-Extended to get out of AI Overviews. It doesn't: Google says Google-Extended doesn't affect Search, and AI Overviews are built from the regular Googlebot index. The only lever there is Googlebot itself (or nosnippet), which also affects normal search results.
- Expecting robots.txt to stop user-triggered fetches. OpenAI says robots.txt may not apply to ChatGPT-User; Perplexity says Perplexity-User generally ignores it; Meta says the same of Meta-ExternalFetcher.
- Ending a group with a bare User-agent line. A User-agent line that comes after a rule starts a new group, so "User-agent: GPTBot, Disallow: /, User-agent: CCBot" with nothing after it leaves CCBot with no rules — and unblocked.
- Blocking the bot but not checking that it stopped. Crawlers cache robots.txt (Google for up to 24 hours; Amazon and Meta say changes can take about 24 hours; DuckDuckGo up to 72) — and some ignore it. Verify in your logs.
- Trusting the user-agent string. Anyone can claim to be GPTBot. Real crawlers come from the vendor's published IP ranges and pass a reverse-DNS check.
How to see which AI crawlers actually visit your site
AI crawlers don't run JavaScript, so script-based analytics — GA4, Plausible, and VisitTrack's own tracker — never see them. You need server-side data: your CDN or server logs, or a few lines of middleware that forward each request's user-agent. VisitTrack's AI crawler tracking does the latter: it classifies each hit as AI answers, indexing or training, verifies the big crawlers by reverse DNS, and shows which pages ChatGPT, Claude and Perplexity fetch most. Pair it with an llms.txt file to point those bots at your best pages.
Want to check a single user-agent string from your logs? The user-agent bot checker tells you whether it's a training crawler, an AI search bot, a search engine, an SEO tool or a human browser.
Frequently asked questions
How do I block GPTBot in robots.txt?
Add a group with "User-agent: GPTBot" followed by "Disallow: /". That stops OpenAI's training crawler. It doesn't affect ChatGPT search (OAI-SearchBot) or pages fetched live for a user (ChatGPT-User), which have their own tokens.
Does blocking AI crawlers remove my site from ChatGPT or Perplexity answers?
Only if you block their search bots. Blocking training crawlers like GPTBot or ClaudeBot doesn't remove you from answers; blocking OAI-SearchBot, Claude-SearchBot or PerplexityBot does remove you from those products' search results and citations over time.
Does Google-Extended block AI Overviews?
No. Google states that Google-Extended controls whether content is used to train Gemini models and to ground Gemini and Vertex AI, and that it has no effect on Google Search inclusion or ranking. AI Overviews and AI Mode use the regular Googlebot index.
Do AI crawlers respect robots.txt?
The major training and search crawlers from OpenAI, Anthropic, Google, Apple, Meta, Amazon, Perplexity and Common Crawl say they do. User-triggered fetchers often don't (OpenAI, Perplexity and Meta say robots.txt may not apply), and ByteDance's Bytespider is widely reported to ignore it. For hard blocking, use your CDN or firewall.
What is the difference between ClaudeBot, Claude-User and Claude-SearchBot?
ClaudeBot collects content that may be used for training Anthropic's models, Claude-SearchBot indexes pages to improve Claude's search results, and Claude-User fetches a page when a Claude user asks for it. Anthropic says all three honor robots.txt, so you can allow or block each independently.
Should I still include anthropic-ai in robots.txt?
It's harmless to keep. anthropic-ai is an older token that Anthropic no longer lists; its current crawlers are ClaudeBot, Claude-SearchBot and Claude-User. The generator includes it as a legacy entry for completeness.
How long does a robots.txt change take to apply?
Usually within a day. Google generally caches robots.txt for up to 24 hours, Amazon and Meta say about 24 hours, and DuckDuckGo says up to 72 hours for DuckAssistBot. Pages already collected for training aren't removed by a later block.
Related tools and guides
- robots.txt testerTest whether a URL is allowed for a crawler, with the matching rule.
- llms.txt generatorPoint AI assistants at the pages that explain your product.
- User-agent bot checkerClassify a UA string as AI crawler, search bot or human.
- Track AI crawlersSee which pages ChatGPT, Claude and Perplexity fetch.
- Bot filteringHow VisitTrack separates humans from bots.
- Analytics glossaryDefinitions for crawler, bot and attribution terms.