Grounding is what separates an AI answer written from a language model’s memory from one that rests on verifiable sources. When someone asks ChatGPT, Gemini, Claude, or Perplexity about a local plumbing company’s business hours, for example, the system can draw on the company’s current website and cite it in the answer instead of relying only on what its model learned during training. Google and Microsoft, among others, use the term in their documentation for Gemini and Copilot.
How grounding works
Without grounding, a language model answers from its parametric knowledge—what it stored during training. That knowledge cannot reliably be traced back to a specific source. With grounding, the model also receives relevant content before it writes and bases its answer on that content. Google describes grounding as the ability to connect a model’s output to verifiable sources of information.
How the content reaches the model depends on the method behind it, usually retrieval-augmented generation (RAG): the system retrieves relevant content—for answers that draw on the web, typically from a web index, a collection of web pages that crawlers (programs that fetch web pages automatically) have captured in advance. In its guide to the AI features in Google Search, Google treats RAG and grounding as two names for the same technique. Where the two are distinguished, as in Microsoft’s documentation, RAG describes the path: retrieve content first, then write the answer. Grounding describes the goal: an answer whose statements can be traced back to specific sources. That is why grounding also applies when the system does not choose the sources itself but is given them, for example web addresses passed to the model along with the question.
Grounding usually becomes visible through source references. Microsoft states that when Copilot’s answers are based on web content, the websites used are listed as links below the text. For Gemini and Claude, the programming interfaces (APIs) that developers use even assign each source reference to a specific passage of the answer. This makes it possible to check what a statement is based on.
Why grounding matters for GEO
Grounding is how a website becomes a source for an AI answer: a system can only reliably give a page as a source—an AI citation—if it draws on that page for the answer. When a model names a business only from what it learned in training, it usually gives no source; that is a mention. In Bing’s Webmaster Guidelines, Microsoft describes GEO as work on making content eligible for grounding and for reference in AI answers. Grounding can also bring information into an answer that only emerged after a model’s knowledge cutoff, such as a new product or changed prices.
However, a system can only draw on what it can reach. According to Google, the AI features in Google Search use publicly accessible, crawlable content for their answers. Important pages should therefore be reachable by crawlers, included in the web index of the system in question, and deliver their main content in the HTML the server sends rather than loading it later with JavaScript (JavaScript rendering). How well crawlers can reach and read a website’s pages is described by its crawlability.
Gemini adds a special case: whether Gemini Apps may use a website’s content for grounding is controlled by the rule for Google-Extended in the site’s robots.txt, the file that sets rules for crawlers. The same rule also decides whether Google may use the content to train future Gemini models.
What grounding cannot do
Grounding makes answers more reliable, but not error-free. Google states that grounding reduces hallucinations and also names several points where errors can arise: in the query the system uses to retrieve sources, in how it interprets the results, and in how it turns them into the answer. Even an answer with a source reference can therefore contain a hallucination.
Moreover, not every answer is grounded: AI systems usually decide for each question whether to retrieve content; according to Anthropic, Claude, for example, does so when a question depends on current or changing information, such as details about specific organizations or products.