When someone asks ChatGPT, Gemini, Claude, or Perplexity for a good provider in their area or for a comparison of two products, a large language model (LLM) writes the answer, including the companies that appear in it. To understand how an offering shows up in AI answers, it helps to know how an LLM is made and how it writes.
What makes a language model large
A language model calculates how likely each possible next piece of text is. These pieces are called tokens; a token is often a word, sometimes just part of a word or a single character. A language model’s parameters are often billions of numerical values that are adjusted step by step during training so that the model gets better and better at predicting how a text continues.
“Large” refers to the number of these parameters and to the amount of text a model learns from; there is no fixed threshold. For a sense of scale: GPT-3, which OpenAI presented in May 2020, had 175 billion parameters. For later models, providers often no longer publish such figures; OpenAI, for example, withheld the model size of GPT-4 in March 2023.
How an LLM is made
An LLM is created in two steps:
- Pretraining: The model processes vast amounts of text and learns to predict the next token each time. Along the way, it picks up how language works, how words and ideas relate, and many facts, including facts about companies and products. These texts are part of its training data; what stays stored in the model is called parametric knowledge and only extends to a certain date, its knowledge cutoff.
- Fine-tuning: A model that has only been pretrained continues text but is not inherently good at answering questions. Providers therefore train it further, using, among other things, example answers written by people and ratings in which people rank several of the model’s answers by quality. Training with such ratings is called reinforcement learning from human feedback (RLHF). Only then does the model answer questions like an assistant, follow instructions, and stick to its provider’s rules.
Fine-tuning shapes how a model answers: how detailed it is, how cautious, and whether it openly acknowledges uncertainty, including in answers about companies and products. Because each provider trains its models with its own data and rules, the same question can get different answers in different AI systems.
How an LLM writes an answer
The input an LLM responds to is called a prompt. From it, the model generates its answer token by token: it calculates which continuation is likely to fit, adds one token, and repeats this until the answer is complete. An LLM does not look anything up in a database; it writes every answer anew. That is why it can fluently and convincingly state something untrue, which is called a hallucination. Because it also has some leeway in choosing each continuation, the same question can lead to different answers, known as answer variability.
An LLM on its own cannot open web pages; it only knows what it learned in training and what is in its input. So that answers can include current information, AI assistants such as ChatGPT, Gemini, Claude, and Perplexity give the underlying model access to web content: depending on the system, the model decides for itself whether to request such content, or content is retrieved for every question. The software around the model does the actual fetching and passes the content to the model along with the question. This method is called retrieval-augmented generation (RAG); the model then bases its answer on the retrieved sources and can cite them as AI citations.
What this means for your visibility
What an LLM writes about your company comes from its parametric knowledge, from retrieved content, or from both. There are therefore two routes to your AI visibility: what future models learn in training, and what AI systems retrieve when a question is asked and can cite as a source. There is no entry in an LLM that a company could edit directly. A new model version can also change how an AI system portrays an offering, even if the published content has stayed the same.
LLMs and related terms
- Language model: the broader term, which also covers smaller and older models. In everyday use, “language model” usually means an LLM.
- Generative AI: the umbrella term for AI models that create new content such as text, images, audio, or video. LLMs are a form of generative AI that mainly produces text.
- Foundation model: a model trained on broad data at scale that can be adapted to many tasks, such as an LLM or a model that generates images.
- AI assistant: the product people talk to. It is built on one or more LLMs and adds features such as retrieving web content; an AI system that retrieves content this way and writes answers from it is also called a generative engine. For Gemini and Claude, the assistant and the model family share the same name. The model that answers can change: Perplexity, for example, uses different models depending on the question, including models from other providers.