LTLLMTXT.co
Control Token · Google

What Is Google-Extended?

Google-Extended is one of the most misunderstood entries in the AI-crawler conversation, because it is not a crawler at all. It is a robots.txt user-agent token that controls whether your content may be used for Google's generative AI — independent of ordinary Google Search indexing done by Googlebot.

https://

A control token, not a bot

When Google crawls your site for Search, it uses Googlebot. Google-Extended does not send its own requests and will not appear as a separate user agent hitting your server. Instead, it is a token you place in robots.txt to express a permission: whether the content Google already accesses may be used to train and ground its generative AI products, which Google has associated with Gemini and Vertex AI. Think of it as a use-policy switch layered on top of Google's existing access, rather than a new fetcher.

This distinction matters because a lot of advice online incorrectly lumps Google-Extended in with true crawlers. Understanding that it is a control token clears up the most common confusion — namely, whether blocking it affects your search rankings.

Google-Extended vs. Googlebot

These two are entirely separate levers, and conflating them leads to costly mistakes:

Googlebot

The crawler that indexes your pages for Google Search. Blocking Googlebot can remove you from search results — a serious decision with major traffic implications.

Google-Extended

A robots.txt token controlling generative-AI training and grounding use. Blocking it has no effect on Search indexing; it only opts your content out of Google's generative-AI use.

In short: you can opt out of generative-AI use with Google-Extended while remaining fully indexed in Google Search. The two decisions are independent.

How to opt out

If you want to keep your content out of Google's generative AI training and grounding while staying in Search, add this to robots.txt:

# Opt out of Google generative-AI use, keep Search indexing
User-agent: Google-Extended
Disallow: /

# Googlebot is untouched, so Search indexing continues
User-agent: Googlebot
Allow: /

Because Google-Extended is only a permission signal, the direction of the tradeoff is different from blocking a crawler: you retain search traffic but reduce the chance your content informs Google's generative answers. As with any robots.txt directive, this is honored by policy rather than technically enforced.

Weighing the tradeoff

Opting out with Google-Extended is attractive for publishers who want their content in Search but not used to train generative models. The cost is potential visibility: if Google's generative surfaces rely on grounded content and yours is excluded, you may be represented less often or less accurately there. Businesses focused on AI discovery often leave Google-Extended allowed, while rights-sensitive publishers frequently opt out. Present the choice neutrally to stakeholders — it is a values-and-strategy decision, not a purely technical one.

Frequently Asked Questions

Is Google-Extended a crawler?

No. Google-Extended is a robots.txt user-agent token, not a distinct bot that fetches your pages. You will not see Google-Extended making requests in your logs the way you see Googlebot. It only controls how content Google already accesses may be used for generative AI.

Does blocking Google-Extended hurt my Google Search ranking?

No. Google-Extended is independent of Googlebot and Search indexing. Disallowing Google-Extended controls generative-AI training and grounding use only; your normal Search presence, governed by Googlebot, is unaffected.

What exactly does Google-Extended control?

Google documents Google-Extended as controlling whether your content helps improve and ground Google's generative AI products, such as Gemini and Vertex AI grounding. It is a use-permission signal, not a page-access control.

How do I opt out?

Add a robots.txt rule for User-agent: Google-Extended with Disallow: /. This signals that your content should not be used for Google's generative AI training and grounding, while leaving Search indexing intact.

Is honoring Google-Extended guaranteed?

Google documents that it respects the directive. As with all robots.txt controls, it is a policy Google chooses to honor rather than a technically enforced block, and the ecosystem continues to evolve.

Understand your AI footprint

Generate an llms.txt file and see how AI systems, including Google's, can find and use your content.

https://