LTLLMTXT.co
Crawler Explainer · OpenAI

What Is GPTBot? OpenAI's Web Crawler

GPTBot is the web crawler operated by OpenAI. Its primary documented role is to collect publicly available web content that can be used to help train OpenAI's models. If you have seen GPTBotin your server logs, this page explains what it is, how it differs from OpenAI's other agents, and how to allow or disallow it.

https://

What GPTBot does

GPTBot requests pages from your website the way any crawler does, following links and reading content. OpenAI describes GPTBot as the agent used to crawl web pages that may be used to improve future models. Because it is a genuine fetcher, GPTBot has its own user-agent string and OpenAI publishes IP ranges you can use to verify that traffic claiming to be GPTBot is authentic.

Allowing GPTBot means your public content can contribute to how models understand your domain, products, and expertise. Disallowing it signals that you do not want your content collected for that purpose. As of mid-2026, OpenAI documents that GPTBot honors robots.txt directives, so a correctly written rule is the standard way to express your preference.

GPTBot vs. OpenAI's other agents

A common point of confusion is treating “OpenAI” as a single crawler. In reality OpenAI operates several distinct agents, each with its own token and purpose:

GPTBot

The training-oriented crawler. This is the agent most people mean when they talk about “blocking OpenAI.”

OAI-SearchBot

Associated with ChatGPT's search experience — surfacing and linking to sources. Controlling this token is distinct from controlling GPTBot, and the two decisions can be made independently.

ChatGPT-User

A user-triggered fetcher — it retrieves a page because a person asked ChatGPT to look at it, rather than as part of a bulk crawl. This is closer to an on-demand browse than a training sweep.

Because these are separate tokens, you might, for example, allow OAI-SearchBot so ChatGPT can cite you while disallowing GPTBot to opt out of training. Confirm the exact current token names in OpenAI's official documentation before writing rules.

How to allow or block GPTBot

Control happens in your robots.txt file at your site root. To disallow GPTBot entirely:

User-agent: GPTBot
Disallow: /

To explicitly allow it (the default behavior if no rule exists), or to allow it while restricting a private section:

User-agent: GPTBot
Allow: /
Disallow: /members/

Remember that robots.txt is advisory. It works because reputable crawlers choose to obey it, not because it can force compliance. If you need a hard guarantee, enforce access at the server or firewall level. And note that llms.txt does not block GPTBot — it is a separate, advisory file for curating content, not an access-control mechanism.

The visibility tradeoff

Deciding whether to allow GPTBot is a strategic choice, not just a technical one. Allowing it can help AI systems represent your business accurately when users ask about your industry. Blocking it protects your content from being used for training but may, over time, reduce how often and how accurately you appear in AI-generated responses. There is no universally correct answer — publishers with proprietary or paywalled content often restrict crawlers, while businesses seeking discovery frequently allow them.

Frequently Asked Questions

What is the GPTBot user agent string?

The robots.txt token is GPTBot. OpenAI publishes the full user-agent string and supporting IP ranges in its official crawler documentation, which is the authoritative source to verify against.

Is GPTBot the same as ChatGPT?

No. GPTBot is the crawler that gathers public web content. ChatGPT is the product. OpenAI runs separate agents — OAI-SearchBot for ChatGPT search and ChatGPT-User for user-triggered fetches — that serve different purposes than GPTBot.

Will blocking GPTBot remove me from ChatGPT search?

Not necessarily. GPTBot is associated primarily with training. ChatGPT's browsing and search features use OAI-SearchBot and ChatGPT-User, which have their own tokens. Blocking one does not automatically block the others.

Does GPTBot respect robots.txt?

OpenAI documents that GPTBot respects robots.txt directives. That said, robots.txt is advisory and cannot be technically enforced, so treat it as a strong signal to a cooperating crawler rather than a guaranteed block.

Should I block GPTBot?

It depends on your goals. Blocking it can keep your content out of future training runs, but it may also reduce your presence in AI-generated answers over time. Weigh the visibility tradeoff against your content-control priorities.

Check how AI reads your site

Generate an llms.txt file and see how discoverable your content is to AI answer engines like ChatGPT.

https://