Retrieval-augmented generation (RAG) follows a simple principle: look things up first, then write. Without RAG, a large language model answers a question from memory alone—from the parametric knowledge it stored during training. With RAG, it is handed texts that match the question and bases its answer on them. AI systems such as ChatGPT, Gemini, Claude, and Perplexity use this principle when they retrieve web content for an answer. For businesses, RAG is therefore the path through which current content from their website can make it into AI answers and appear there as a source.

How RAG builds an answer

The name describes the method’s three steps. For example, someone asks an AI system which hotels in a particular city offer meeting rooms for 80 people.

  1. Retrieval: The system looks through a collection of documents for passages that match the question. For AI assistants, this is usually a web index: a database of web pages that crawlers—programs that fetch web pages automatically—have captured in advance. The system often derives several subqueries from the question, for example on conference hotels, room sizes, and location; this approach is called query fan-out.
  2. Augmentation: The passages found are passed to the language model together with the original question. They augment the prompt, the input the model responds to.
  3. Generation: The language model writes the answer based on the question and the passages it received—in the example, from the pages of hotels and event venue portals. Many AI systems point to the pages they used. These source references are called AI citations.

The name comes from a paper by researchers at Facebook AI Research, University College London, and New York University, presented at the NeurIPS conference in 2020. Their system searched around 21 million short passages from Wikipedia. The researchers argued that language models that store their knowledge only within themselves cannot easily expand or revise it, make it hard to see what an answer is based on, and sometimes make up facts. Retrieved texts, by contrast, can be read by humans and swapped out. In one experiment, the researchers ran their system once with a December 2016 version of Wikipedia and once with a December 2018 version. It mostly answered questions such as “Who is the President of Peru?” according to the version in use—without being retrained.

What RAG means for your website’s visibility

Because RAG draws on content at the moment a question is asked, current information from a website can flow into answers once the index has picked up the current version of the page. What a model knows without retrieval, by contrast, only goes up to its knowledge cutoff. To appear in an answer through RAG, however, a page has to pass several stages. The following breakdown is a simplification, since every system works differently in detail:

  • Being captured: As a rule, a crawler must have fetched the page before anyone asks, and it must be in the index of the system in question. Whether crawlers are allowed and able to fetch a page is what crawlability is about.
  • Being found: During retrieval, the page or a section of it must match one of the queries the system derives from the question.
  • Being selected: Many RAG systems assess the passages found more closely in a second step (reranking) and pass only the most relevant ones to the language model.
  • Being used: In the last step, the system determines which content goes into the answer and which pages appear as sources. Sections that make sense on their own are easier to use; this property is called citability.

Each stage acts as a filter: what has not been captured cannot be found, and a passage that is not selected never reaches the language model. A study presented at the KDD 2026 conference examined how much depends on the early stages, in a test setup that simulates retrieval, selection, and answer generation. Texts whose body copy alone had been rewritten using methods from GEO research gained hardly any visibility at the answer stage and dropped out more often during retrieval and selection. When the rewrites also covered the page title, the short summary in the page’s code, the headings, and structured data, they performed better, especially during retrieval. However, the test setup evaluates these elements explicitly and does not replicate any provider’s systems. What it does show is that the effect of a change can only be judged across all stages.

RAG and grounding

A closely related term is grounding. In its guide to the generative AI features in Google Search, Google treats the two as the same: it describes RAG as a technique, “also known as grounding,” that improves the quality, accuracy, and freshness of AI responses by retrieving relevant, up-to-date web pages from Google’s index. Where the two are distinguished, RAG refers to the method—retrieve first, then write—and grounding to the goal: an answer whose statements can be traced back to specific sources.

Limits of RAG

RAG can only work with what retrieval delivers. Anthropic, the company behind Claude, points out that the effectiveness of RAG depends on the quality and relevance of the underlying collection and of the content retrieved. Retrieval can miss a relevant page or return a passage that only partly answers the question. A passage taken out of its page can also lose its context, for example when it does not name the company it is about. An answer based on retrieved content is therefore not automatically free of errors and can even contain hallucinations.

RAG also does not come into play for every answer. Whether an AI system retrieves content for a question at all depends on the system, its settings, and the question; when it answers from its parametric knowledge alone, RAG plays no part in that answer. Providers document only in broad terms how they select, weigh, and combine content.