How to Block AI Crawlers
If you want to keep your content out of AI training data and answer engines, the documented mechanism is a set of robots.txt rules that name each AI crawler and disallow it. This guide shows exactly how — and explains the visibility tradeoff so you can decide with eyes open.
The copy-paste block
Place the following in your robots.txt file at your domain root (for example, https://yourdomain.com/robots.txt). It disallows the most commonly discussed AI crawlers and control tokens while leaving search crawlers untouched:
# Block major AI training / answer-engine agents User-agent: GPTBot Disallow: / User-agent: OAI-SearchBot Disallow: / User-agent: ChatGPT-User Disallow: / User-agent: ClaudeBot Disallow: / User-agent: anthropic-ai Disallow: / User-agent: PerplexityBot Disallow: / User-agent: Bytespider Disallow: / # Control tokens (training-use opt-out, not crawlers) User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: /
Token names evolve as of mid-2026, so verify each against the operator's official documentation and add any new agents that appear. Removing a block for a given agent is as simple as deleting its two lines.
Understand what each line does
Not every token in that list behaves the same way, and knowing the difference prevents mistakes:
True crawlers
GPTBot, ClaudeBot, PerplexityBot, and Bytespider actually fetch your pages. Disallowing them asks them to stop requesting your content. You can confirm the effect by watching your access logs.
Control tokens
Google-Extended and Applebot-Extended are not separate fetchers. Disallowing them opts your content out of Google's and Apple's generative-AI training use, without affecting their base search crawlers.
Do not block Googlebot by accident
Keep AI tokens separate from Googlebot and Bingbot. Disallowing those search crawlers can remove you from search results — a very different and much larger consequence.
When robots.txt is not enough
robots.txt works because cooperating crawlers choose to obey it. It has no teeth against a crawler that ignores it or spoofs a different user agent. If you need a genuine, enforceable block — for example, against a crawler you have disallowed but still see in your logs — implement it at the server or firewall level, where you can reject requests outright:
# Apache example: deny selected AI user agents <IfModule mod_setenvif.c> BrowserMatchNoCase "GPTBot" ai_bot BrowserMatchNoCase "ClaudeBot" ai_bot BrowserMatchNoCase "Bytespider" ai_bot </IfModule> <RequireAll> Require all granted Require not env ai_bot </RequireAll>
Server-level rules are enforceable, but they also carry more risk of misconfiguration, so test carefully to avoid blocking legitimate visitors.
The tradeoff, stated plainly
Blocking AI crawlers is a legitimate choice, but it is not free. As AI answer engines become a larger part of how people discover businesses and information, excluding your content from those systems can reduce how often you are mentioned, cited, or accurately described. If your priority is protecting proprietary or paid content, blocking makes sense. If your priority is discovery, allowing selected crawlers — and helping them parse your content — is usually the better path. We present this neutrally: the right answer depends on your business, not on a one-size-fits-all rule.
Frequently Asked Questions
Does robots.txt actually stop AI crawlers?
For reputable crawlers that document respect for robots.txt, yes — they will honor a correctly written rule. But robots.txt is advisory and cannot technically enforce anything. A misbehaving or unofficial crawler can ignore it, so use server-level blocking when you need a hard guarantee.
Will blocking AI crawlers hurt my Google Search ranking?
Blocking AI-specific tokens like GPTBot, Google-Extended, or Applebot-Extended does not affect Google Search, which is governed by Googlebot. Just be careful not to accidentally disallow Googlebot itself, which would affect search indexing.
What is the downside of blocking AI crawlers?
Blocking training and answer-engine crawlers can reduce how often and how accurately your business appears in AI-generated answers over time. That is the core tradeoff: more content control, potentially less AI visibility.
Can I block AI crawlers but keep search crawlers?
Yes. Because each crawler has its own user-agent token, you can disallow AI crawlers while allowing Googlebot, Bingbot, and other search crawlers. This is a common configuration for publishers who want search traffic without AI training use.
Does llms.txt block crawlers?
No. llms.txt is advisory content curation for language models. It does not block anything. Blocking happens in robots.txt or at the server level.
Not sure whether to block or allow?
Check how AI systems currently see your site, then decide. Enter your domain to get started.