LTLLMTXT.co
How-To Guide

How to Block AI Crawlers

If you want to keep your content out of AI training data and answer engines, the documented mechanism is a set of robots.txt rules that name each AI crawler and disallow it. This guide shows exactly how — and explains the visibility tradeoff so you can decide with eyes open.

https://

The copy-paste block

Place the following in your robots.txt file at your domain root (for example, https://yourdomain.com/robots.txt). It disallows the most commonly discussed AI crawlers and control tokens while leaving search crawlers untouched:

# Block major AI training / answer-engine agents
User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: anthropic-ai
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: Bytespider
Disallow: /

# Control tokens (training-use opt-out, not crawlers)
User-agent: Google-Extended
Disallow: /

User-agent: Applebot-Extended
Disallow: /

Token names evolve as of mid-2026, so verify each against the operator's official documentation and add any new agents that appear. Removing a block for a given agent is as simple as deleting its two lines.

Understand what each line does

Not every token in that list behaves the same way, and knowing the difference prevents mistakes:

True crawlers

GPTBot, ClaudeBot, PerplexityBot, and Bytespider actually fetch your pages. Disallowing them asks them to stop requesting your content. You can confirm the effect by watching your access logs.

Control tokens

Google-Extended and Applebot-Extended are not separate fetchers. Disallowing them opts your content out of Google's and Apple's generative-AI training use, without affecting their base search crawlers.

Do not block Googlebot by accident

Keep AI tokens separate from Googlebot and Bingbot. Disallowing those search crawlers can remove you from search results — a very different and much larger consequence.

When robots.txt is not enough

robots.txt works because cooperating crawlers choose to obey it. It has no teeth against a crawler that ignores it or spoofs a different user agent. If you need a genuine, enforceable block — for example, against a crawler you have disallowed but still see in your logs — implement it at the server or firewall level, where you can reject requests outright:

# Apache example: deny selected AI user agents
<IfModule mod_setenvif.c>
  BrowserMatchNoCase "GPTBot" ai_bot
  BrowserMatchNoCase "ClaudeBot" ai_bot
  BrowserMatchNoCase "Bytespider" ai_bot
</IfModule>
<RequireAll>
  Require all granted
  Require not env ai_bot
</RequireAll>

Server-level rules are enforceable, but they also carry more risk of misconfiguration, so test carefully to avoid blocking legitimate visitors.

The tradeoff, stated plainly

Blocking AI crawlers is a legitimate choice, but it is not free. As AI answer engines become a larger part of how people discover businesses and information, excluding your content from those systems can reduce how often you are mentioned, cited, or accurately described. If your priority is protecting proprietary or paid content, blocking makes sense. If your priority is discovery, allowing selected crawlers — and helping them parse your content — is usually the better path. We present this neutrally: the right answer depends on your business, not on a one-size-fits-all rule.

Frequently Asked Questions

Does robots.txt actually stop AI crawlers?

For reputable crawlers that document respect for robots.txt, yes — they will honor a correctly written rule. But robots.txt is advisory and cannot technically enforce anything. A misbehaving or unofficial crawler can ignore it, so use server-level blocking when you need a hard guarantee.

Will blocking AI crawlers hurt my Google Search ranking?

Blocking AI-specific tokens like GPTBot, Google-Extended, or Applebot-Extended does not affect Google Search, which is governed by Googlebot. Just be careful not to accidentally disallow Googlebot itself, which would affect search indexing.

What is the downside of blocking AI crawlers?

Blocking training and answer-engine crawlers can reduce how often and how accurately your business appears in AI-generated answers over time. That is the core tradeoff: more content control, potentially less AI visibility.

Can I block AI crawlers but keep search crawlers?

Yes. Because each crawler has its own user-agent token, you can disallow AI crawlers while allowing Googlebot, Bingbot, and other search crawlers. This is a common configuration for publishers who want search traffic without AI training use.

Does llms.txt block crawlers?

No. llms.txt is advisory content curation for language models. It does not block anything. Blocking happens in robots.txt or at the server level.

Not sure whether to block or allow?

Check how AI systems currently see your site, then decide. Enter your domain to get started.

https://