Google-Extended is a name a website can address in its robots.txt file to decide whether its content may be used for Gemini training and grounding. Google introduced it in September 2023, at the time for its AI assistant Bard, which has been called Gemini since February 2024, and for the generative AI APIs (programming interfaces for developers) of its cloud platform Vertex AI. Two points are easily misunderstood: Google-Extended looks like the name of a crawler, a program that fetches web pages automatically, but it isn’t one. And the rule does not decide whether a website appears in the AI features of Google Search.

A token without a crawler

In robots.txt, rules address crawlers by their token, the name a crawler identifies itself with. Google-Extended is such a token, but according to Google there is no crawler of its own behind it: Google crawls with its existing crawlers and uses the token only as a control. A rule for Google-Extended therefore doesn’t decide whether Google fetches a page, but what Google may use the fetched content for.

The rule goes into a group of its own in robots.txt. The lines “User-agent: Google-Extended” and “Disallow: /” exclude the content of the entire website from the uses that Google-Extended governs. As with any other group, the rule can also cover only certain directories, such as an archive.

What Google-Extended controls

According to Google, the token governs two uses:

  • Training: whether content may be used to train future generations of Gemini models, that is, whether it may become part of their training data. Google names the models that power Gemini Apps and the Vertex AI API for Gemini.
  • Grounding: whether Gemini may use the content for grounding. Google defines this as providing content from Google’s web index to the model at the moment a question is asked, to make the answer more factual and relevant. This applies to grounding in Gemini Apps and to the Grounding with Google Search tool on Gemini Enterprise Agent Platform, Google’s cloud platform for businesses, which grew out of Vertex AI in April 2026.

What Google-Extended does not control

According to Google, Google-Extended does not affect whether a website is included in Google Search. That also applies to AI Overviews and AI Mode, which are part of Google Search: whether content may appear there is controlled through access for Google’s crawler Googlebot, directives in a page’s code such as nosnippet, and the “Search generative AI” setting in Google Search Console, Google’s tool for website owners. To limit training of the models behind these features, however, Google points to Google-Extended.

The rule also refers to training future generations of Gemini models; Google’s documentation on Google-Extended describes no process for removing content from models that have already been trained. And it concerns content that Google crawls. Fetches that a person triggers, for example by adding a web address as a source in Gemini Notebook, are handled by separate user-triggered fetchers, which, according to Google, generally ignore robots.txt rules. Google does not say whether Google-Extended applies to content from such fetches.

Training and grounding go together

For an AI training opt-out at Google, Google-Extended is the route Google documents. With OpenAI and Anthropic, training and use as a source for answers can be controlled through separate tokens—at OpenAI, for example, with GPTBot for training. With Gemini, both depend on the same rule. Anyone who blocks Google-Extended to keep content out of training is making a second decision at the same time: according to Google, Gemini Apps may then no longer use the website’s pages to ground their answers.

For Gemini, that leaves a single trade-off. Some businesses want Gemini Apps to be able to use their current pages for grounding and allow Google-Extended; in doing so, they also allow their content to be used to train future Gemini models. Others deliberately opt out of training and give up grounding in return. Whether a Google-Extended rule also counts as a TDM opt-out in the legal sense is a separate question of copyright law.