Prompt tracking is the step between choosing the questions and calculating the metrics. Which questions are asked is set by a prompt set, a defined collection of typical customer questions; prompt tracking is where those questions are actually asked and the answers collected, in large numbers. For example, 50 prompts, each asked seven times a day in ChatGPT, Gemini, Claude, and Perplexity, add up to 1,400 answers a day. Volumes like these can only be handled by software that asks the questions on a schedule.
Why a single query is not enough
AI systems don’t answer the same question the same way every time. One answer names a brand, the next one doesn’t, and the cited sources change as well. These fluctuations are known as answer variability. For measurement, this means that only many repetitions of the same question show how often a brand actually comes up.
How many repetitions are needed is not settled. In an April 2026 preprint, a study released ahead of peer review, researchers at the University of St. Gallen recommended asking each question at least seven times a day and combining the results over two to four weeks. Their data came from German-language questions sent to ChatGPT, Gemini, Google’s AI Mode, and Perplexity from servers in Switzerland between January and March 2026. A September 2026 preprint reanalyzed five earlier studies by its author’s own research group on brand recommendations by language models and derived tiers: 5, 10, or 15 repetitions, depending on how reliable the result needs to be. It also found, however, that these fixed tiers did not carry over to three independent datasets; its author sees this as a reason to run a pilot on the questions that matter before settling on a number. The first author of the St. Gallen study and the author of the second paper are both also affiliated with companies that sell AI visibility measurement.
What is fixed and recorded
Results are only comparable if, between two measurements, only the answers change—not the conditions under which they are produced. That is why the settings are fixed once and stored with every answer:
- AI systems: which systems are queried, for example ChatGPT, Gemini, Claude, and Perplexity, and, where there is a choice, which model and whether web search is turned on.
- Region: which country the questions are sent from. Some AI systems infer an approximate location from the IP address, the network address a request comes from, and may take it into account for location-related questions. The St. Gallen researchers point out that their results, collected in Switzerland, may not carry over to other markets.
- Language: which language and country settings apply, matched to the language of the prompts and to each market.
- Access: whether the questions go through the apps’ interfaces, the way people use them, or through an application programming interface (API), which lets software address a language model directly. The two don’t necessarily produce the same answers: through OpenAI’s API, a model only searches the web when web search is set up explicitly, either as a tool or by choosing a dedicated search model. And according to Anthropic, Claude’s web interface and mobile apps give the model additional instructions that do not apply when it is accessed through the API. The St. Gallen study queried the interfaces; the five reanalyzed studies used APIs.
- Conversation state: each question in a new conversation, without earlier chats, saved memories, or custom instructions, since some AI systems, such as Gemini, take these into account depending on the settings.
- Time and version: the date and time of every answer and, where shown, the model version. An unchanged version name does not rule out silent changes, though: providers can adjust their systems without changing the version, and Anthropic periodically updates the instructions Claude receives in its web interface and mobile apps.
If the setup changes, for example because an AI system is added, the change is recorded with its date. That way, a jump in the results can later be checked against changes to the setup.
Recording and analyzing answers
Each answer is stored in full: the answer text, the cited sources with their links, and the settings it was produced under. The analysis then checks, answer by answer, whether the brand or a competitor is named, including variant spellings (a mention), whether one of the brand’s pages appears as a source (an AI citation), and whether the statements about the brand are correct (brand accuracy). If a language model does this analysis, its instructions become part of the fixed setup too: a review paper that examines existing methods for measuring AI visibility, also released as a preprint in September 2026, points out that a different instruction alone can change the score given to an unchanged answer.
The analyzed answers yield metrics such as mention rate, reported separately for each AI system. Only their development over weeks and months shows how a brand’s visibility is changing; observing and reporting it on an ongoing basis is called GEO monitoring.