Do AI Crawlers Respect llms.txt?
The short, honest answer: llms.txtis not something a crawler “respects” or “disobeys,” because it contains nothing to obey. It is an advisory curation file, not an access-control mechanism. Understanding that distinction is the key to setting the right expectations.
Two different files, two different questions
The confusion comes from lumping robots.txt and llms.txt together because both are root-level text files. But they answer different questions. robots.txt answers “may this crawler fetch this path?” — it is the long-standing Robots Exclusion Protocol, and it contains explicit allow/disallow directives that crawlers can honor. llms.txt answers “what content matters most on this site?” — it is a curated Markdown map with no directives at all. There is literally no rule in an llms.txt for a crawler to comply with or ignore.
What reputable crawlers document
Where compliance genuinely applies is robots.txt. The major AI crawler operators publish documentation stating that their bots respect robots.txt directives:
- GPTBot — OpenAI's crawler; documented to honor robots.txt.
- ClaudeBot — Anthropic's crawler; documented to honor robots.txt.
- PerplexityBot — Perplexity's crawler; documented to honor robots.txt.
- Google-Extended — a robots.txt user-agent token Google offers to control use of your content for AI training (not a separate crawler).
- Applebot-Extended — Apple's token for controlling AI training use.
- Bytespider — ByteDance's crawler.
Two honest caveats belong here. First, this compliance is voluntary — a stated policy, not a technical enforcement. Well-behaved crawlers follow robots.txt; a determined or poorly-behaved bot can ignore it, and not every crawler on the internet is reputable. Second, all of this is about robots.txt. None of it means a crawler does anything special with your llms.txt.
So what does happen with llms.txt?
At best, a system that supports the convention fetches your /llms.txt, reads your curated summary and links, and uses that as one input into how it understands your site. Whether any given system does this — and how much weight it gives the file — varies across platforms and is not guaranteed. There is no standard body certifying support, no compliance badge, and no way to compel ingestion. As of mid-2026, adoption is still emerging, and the most accurate framing is: publishing llms.txt makes your intent available to systems that choose to look, and does nothing to systems that do not.
That is not a reason to skip it. It is a low-cost, low-risk way to present your content clearly. But it is a reason to keep your mental model accurate: llms.txt is an offer, robots.txt is a request that reputable crawlers honor, and neither is a lock.
If your goal is actually to block AI
Do not reach for llms.txt — it cannot help. Use robots.txt to disallow specific AI crawler user-agents, and back it with server-level rules if you need something stronger than a voluntary convention. A minimal robots.txt that discourages the major AI training crawlers looks like this:
User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: /
Even here, remember the caveat: this relies on each crawler honoring robots.txt, which reputable ones document but none can be forced to do. For a fuller walkthrough, see the guide on blocking AI crawlers.
Frequently Asked Questions
Does llms.txt block crawlers?
No. llms.txt has no defined allow or disallow directives and no enforcement mechanism. It cannot block, throttle, or gate any crawler. If you want to control access, that is a robots.txt or server-configuration job.
Do GPTBot, ClaudeBot, and PerplexityBot obey robots.txt?
Their operators — OpenAI, Anthropic, and Perplexity respectively — publicly document that these crawlers respect robots.txt directives. That is a stated policy, honored voluntarily; it is not a technical guarantee, and not every bot on the internet behaves the same way.
Will publishing llms.txt make AI systems read it?
Not necessarily. Whether a given system fetches and ingests your llms.txt varies by platform and is not guaranteed. Adoption is still emerging as of mid-2026. Publishing it is a low-cost way to make your intent legible, not a switch that forces consumption.
Is llms.txt a privacy or security tool?
No. It is public and advisory. Never treat it as a way to hide content. Anything you would not want quoted should not be linked from it, and genuinely private material must be protected by real access controls.
Can I tell if a crawler used my llms.txt?
Rarely with certainty. Server logs may show requests to /llms.txt from identifiable crawler user-agents, but many systems do not announce themselves and ingestion is not observable from your side. You can confirm the file was fetched, not that it shaped an answer.
Present your content clearly to AI
You can't force ingestion, but you can make your site easy to understand. Generate a clean llms.txt for your domain.