Answer variability often shows up as soon as you ask an AI system the same question twice. For example, if you ask “Which payroll providers are a good fit for mid-sized companies?” the system names five providers the first time, including yours. The second time, it names four, and your company is missing. This is not a malfunction but a consequence of how AI systems such as ChatGPT, Gemini, Claude, and Perplexity produce their answers.

Where the differences come from

Researchers who study these fluctuations distinguish several causes, which often work together.

Chance in word choice

A large language model writes an answer word by word. At each step, it calculates which continuations are likely and picks one of them, with an element of chance. The amount of chance is controlled by a setting called temperature: at a high temperature, answers turn out more varied; at a low one, they tend to stick to the most probable wording. Because every word builds on the ones before it, a single different choice can change the rest of the answer, right down to which providers it lists. The companies behind AI systems decide for themselves which temperature their own apps use.

Changing sources

Many AI systems retrieve current content from the web for an answer, a method called retrieval-augmented generation (RAG). Which pages are found, and which of them an answer ends up citing as sources, can differ from one query to the next. If an AI system splits a question into several subqueries that it writes itself (query fan-out), these can also differ for the same question. Different sources bring different facts, different brands, and different AI citations into the answer.

Different phrasing

People ask the same question in many ways. Even a differently worded prompt, the input given to the AI system, can lead to a different answer. Google, for instance, noted in a May 2025 explainer that its AI Mode may give different answers to two similar questions about the same topic if they are framed differently. How much the sources of an answer change with different wording or a different language also varies from one AI system to the next. That was the finding of a September 2025 preprint, a study released ahead of peer review.

Time and context

Over weeks and months, further causes come into play: providers update their models, and content on the web changes. Some AI systems also take into account the approximate location or earlier conversations of the person asking. Much of the fluctuation, however, occurs without any such changes: in a study by researchers at the University of St. Gallen, also released as a preprint in April 2026, the cited sources differed about as much between repeated queries on the same day as from one day to the next. The researchers sent German-language questions from servers in Switzerland between January and March 2026; the study’s first author is also affiliated with a company that sells AI visibility measurement.

What answer variability means for visibility

A single AI answer is a sample, not a measurement. A screenshot showing your brand at the top, or missing entirely, captures one of many possible answers. The St. Gallen researchers therefore conclude that AI visibility should be understood as the probability of being mentioned across repeated queries. Whether your brand comes up for a question can thus only be expressed as a share: how many out of many answers name it, across repeated queries and different wordings of the same question. This applies not only to mentions but to everything that makes up your AI visibility, such as the position in a list or the accuracy of the details.

Every such share is an estimate with some uncertainty. Two answers reveal very little: if a brand shows up in one of them, it may in fact be named almost always or only rarely. The more answers go into the analysis, the narrower the range in which the true value likely lies. A meaningful report therefore states not only the share but also how many answers it is based on. Studies give different recommendations on how often a question should be repeated in prompt tracking; there is no generally accepted number. One metric built from repeated answers is the mention rate.

For businesses, answer variability also has a reassuring side: an answer without your brand is not a final verdict. Whether visibility is improving shows in the share across many answers, as tracked over time in GEO monitoring, not in a single query.