Training data shapes what an AI system knows about a company when it answers without retrieving any content from the web. The large language models behind ChatGPT, Gemini, Claude, and Perplexity learn from it, including facts about companies, products, and industries. What stays stored in the model is called parametric knowledge. For example, if trade media, directories, and an industrial pump manufacturer’s own website published a lot about the company before training, a model can later name the company without retrieving anything from the web. If the manufacturer barely appeared in the training data, the model knows little or nothing about it.
What training data consists of
Providers assemble training data from several kinds of sources:
- Web content: publicly accessible pages collected by crawlers, programs that fetch web pages automatically. Some AI providers run their own crawlers for this, such as OpenAI with GPTBot. Many language models have also learned from Common Crawl, a freely available archive of web pages that a nonprofit organization collects on a large scale. According to a Mozilla Foundation report from February 2024, most of the 47 language models it examined from the years 2019 to 2023 were trained on Common Crawl data.
- Selected datasets: books, reference works such as Wikipedia, and datasets that providers purchase or obtain through agreements.
- Data from use: Anthropic, for example, uses conversations with Claude for training when users have opted in or submitted them as feedback, or when they were flagged for violating its usage policy.
- Synthetic data: text generated by other AI models.
Providers do not disclose exactly how this mix is composed for commercial models. On its Transparency Hub, for instance, Anthropic names only the kinds of sources for its Claude models, not their shares or individual websites. Older models are documented in more detail, as are open models whose developers also publish their training data. For GPT-3, for example, which OpenAI presented in May 2020, most of the training data came from a filtered version of Common Crawl, supplemented by another web collection, books, and English-language Wikipedia. Sources the researchers considered higher quality, such as Wikipedia, were used several times during training, while the Common Crawl portion was not even used once in full. On top of this come smaller, purpose-built datasets, such as example answers written by people, which providers use to turn a model, after its first, broad round of training, into an assistant that answers questions.
Not everything a crawler collects ends up in training. Providers clean and filter the text first, as Anthropic states for its Claude models. The open dataset FineWeb, released in April 2024 and based on Common Crawl, shows what this can look like: among other things, its creators used a blocklist to remove pages with adult content, along with text that failed quality rules, for example because it consisted mostly of very short or repeated lines. A page made up almost entirely of short fragments without complete sentences can fail such rules.
How a company gets into training data
Anything that is publicly available about a company and that crawlers are allowed to collect can get into training data: the company’s own website, but also trade articles, directories, reviews, or Wikipedia, the sources that Off-Page GEO is about. Crawlers generally do not reach pages behind a login, such as pages in a customer portal; according to Anthropic, its crawler ClaudeBot does not access password-protected pages.
A study presented at the machine learning conference ICML in 2023 showed that language models answer factual questions more accurately the more documents related to the question their training data contained. How well a model knows a company therefore likely also depends on how often the company appears in the training data. For companies, this is an argument for making sure accurate, consistent information appears in many places, not just on their own website—which is what brand fact consistency is about. Training data also only extends to a certain point in time; the date up to which a model’s knowledge is reliable is its knowledge cutoff. Only a later model can learn from training what was published after the training data ends.
Deciding on training use
Whether content from its website should flow into future training data is something a company signals in its robots.txt file, with rules for training crawlers such as GPTBot from OpenAI or ClaudeBot from Anthropic, which at these providers do not affect the crawlers AI systems use to retrieve sources for their answers; at Google, by contrast, the rule for Google-Extended also covers whether the Gemini Apps may use content for their answers. The limits of such a rule, such as content that has already been collected or copies on other websites, are covered under AI training opt-out.