AI Crawler Directory: Bots and User Agents
As AI systems increasingly read the open web, a handful of crawlers and robots.txt control tokens now shape whether your content feeds training data, powers answer engines, or stays out entirely. This directory summarizes the major players and how to reference them accurately.
Why an AI crawler directory matters
Traditional search crawlers like Googlebot exist to index pages for a results list. AI crawlers serve a broader set of purposes: some gather text to train large language models, some fetch pages in real time to answer a specific user question, and some index content so an answer engine can cite it. Because these jobs differ, the companies behind them publish separate user-agent tokens so site owners can allow or disallow each use case independently in robots.txt.
As of mid-2026 this ecosystem is still maturing. Operators add, rename, and document agents over time, and the line between “training,” “grounding,” and “user-triggered” fetching is not always crisp. Treat any single snapshot — including this one — as a starting point, and confirm the current token names against each operator's official documentation before you rely on them.
The major AI crawlers and tokens
The list below covers the agents most commonly discussed by publishers. Some are true crawlers with their own IP ranges; others are robots.txt-only control tokens that gate how already-crawled content may be reused.
GPTBot
OpenAIUser-agent: GPTBotOpenAI's crawler used to gather publicly available web content for training its models. OpenAI also runs OAI-SearchBot (for ChatGPT search results) and ChatGPT-User (user-triggered fetches), which are distinct tokens.
ClaudeBot
AnthropicUser-agent: ClaudeBotAnthropic's crawler for collecting web content. Anthropic has referenced both ClaudeBot and an anthropic-ai user agent in its documentation, plus a user-triggered agent for Claude-initiated fetches.
PerplexityBot
PerplexityUser-agent: PerplexityBotPerplexity's crawler used to index pages that its answer engine can cite. Perplexity also references a PerplexityUser agent for fetches triggered directly by a user's question.
Google-Extended
GoogleUser-agent: Google-ExtendedNot a separate crawler. It is a robots.txt user-agent token that controls whether your content is used for Google's generative AI products (such as Gemini and Vertex AI grounding), independent of ordinary Googlebot indexing.
Applebot-Extended
AppleUser-agent: Applebot-ExtendedA robots.txt control token, not a distinct fetcher. It governs whether content Applebot has already crawled may be used to train Apple's generative models. Applebot remains the base crawler for search features.
Bytespider
ByteDance (TikTok)User-agent: BytespiderByteDance's crawler, associated with data collection for its AI products. Some publishers report aggressive crawling; behavior and robots.txt handling should be monitored rather than assumed.
Crawlers vs. control tokens
It is important not to conflate the two categories. GPTBot, ClaudeBot, PerplexityBot, and Bytespider are fetchers that request pages from your server, so you may see them in your access logs. Google-Extended and Applebot-Extended are different: they are robots.txt user-agent tokens that signal how your content may be used for generative AI, but they do not themselves appear as a distinct bot hitting your server. Blocking Google-Extended does not remove your pages from Google Search — Googlebot indexing is governed separately.
How control actually works
The documented mechanism for allowing or disallowing these agents is a robots.txt rule that names the user agent and sets a Disallow or Allow path. Reputable operators publish documentation, and often IP ranges, and state that they respect robots.txt. However, robots.txt is a voluntary standard: it expresses a preference and cannot technically enforce compliance. A poorly behaved or unofficial crawler can ignore it, and there is no guarantee a given AI system will honor every rule.
# Example: disallow common AI training crawlers User-agent: GPTBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: /
Note that llms.txt is a different tool entirely. It is an advisory file that curates and describes your content for language models; it does not block anything and is not an access-control mechanism.
Frequently Asked Questions
Is Google-Extended a crawler I will see in my logs?
No. Google-Extended is a robots.txt user-agent token that controls whether your content is used for Google's generative AI. It does not fetch pages as a separate bot, and it is independent of Googlebot search indexing.
Does blocking these crawlers remove me from search?
Blocking AI-specific tokens like Google-Extended or Applebot-Extended does not remove you from the corresponding search index, which is governed by Googlebot and Applebot respectively. Blocking a training crawler may, however, reduce your visibility inside AI-generated answers over time.
Can robots.txt guarantee a crawler stays out?
No. robots.txt is advisory. Reputable operators document that they respect it and often publish IP ranges, but compliance cannot be technically enforced by robots.txt alone. Server-level blocking is a stronger control if you need certainty.
How is this different from llms.txt?
llms.txt is advisory guidance that describes and links your key content for language models. It does not allow or block any crawler. robots.txt is where access control lives; llms.txt is where content curation lives.
Do these token names change?
Yes, the ecosystem is still evolving as of mid-2026. Operators occasionally add or rename agents. Always confirm the current names in each company's official crawler documentation before relying on a rule.
See how AI systems read your site
Enter your domain to generate an llms.txt file and check how discoverable your content is to AI answer engines.