Brand accuracy answers a simple question for a brand: Is what AI systems such as ChatGPT, Gemini, Claude, and Perplexity say about it correct? A brand can appear in many answers and still show up with an outdated price, a wrong location, or a service it does not offer at all. Mention rate only shows how often a brand is named; brand accuracy shows whether what is said is true. That makes it a distinct part of a brand’s AI visibility.
How brand accuracy is measured
An AI answer usually contains several claims, and not all of them are necessarily correct. Brand accuracy is therefore assessed statement by statement rather than for the answer as a whole. For example, an AI system is asked, “What does Example Building Services offer?” It replies: “Example Building Services, based in Kassel, was founded in 1995, designs and installs heating systems and bathrooms, and offers a 24-hour emergency service.” Broken down into individual statements and checked against the company’s verified facts, the answer looks like this:
- Based in Kassel: correct.
- Founded in 1995: wrong; the company was founded in 1998.
- Designs and installs heating systems: correct.
- Designs and installs bathrooms: correct.
- 24-hour emergency service: outdated; the service was discontinued in 2022.
Three of the five statements are correct, so the brand accuracy of this answer is 60%. Yet as a whole, the answer sounds plausible, which is exactly what makes such errors hard to spot. Outdated statements count as errors just like wrong ones. It helps, however, to note them separately, because they point to a different cause: sources or training knowledge that still reflect the old state of affairs.
The method is used in research on language models. A well-known example is FActScore, an evaluation method that researchers presented at the EMNLP research conference in 2023. They broke biographies written by language models into the smallest checkable statements, called atomic facts, and checked each one against Wikipedia. The score is the share of statements that the knowledge source supports. For a brand, Wikipedia is replaced by an internal fact sheet with the company’s verified information—the same basis used for brand fact consistency.
Aggregated over many answers and reported separately for each AI system, these checks yield a brand’s overall accuracy figure. It can be expressed as the share of correct statements or as the share of answers without a wrong statement; there is no standard calculation.
Which facts are checked
Checks use the questions customers might ask an AI system about a company: what it offers, what it costs, where it is located, how to reach it, and how long it has been in business. Such questions name the company explicitly. Answers to broader questions are revealing too, for example about providers in a region, when they name the brand along with features or prices. Both kinds of questions belong in a fixed prompt set. Because AI systems do not answer the same question the same way every time (answer variability), the assessment is based on several answers per question.
Not every error is equally serious. A wrong phone number, an outdated price, or a discontinued service sends customers in the wrong direction; a wrong founding year matters less. It therefore makes sense to weight errors by their consequences and tackle the most consequential ones first.
Tracing and correcting errors
An AI system cannot be corrected directly, but the sources it draws on can be. Correction therefore starts with finding out where an error comes from. When an answer gives sources (AI citations), they show what it is probably based on, such as a directory listing with an old address or the website of a company with the same name. An answer can reproduce its source faithfully and still be wrong if the source itself is outdated. Conversely, an answer can also misrepresent a source that is correct. When there is no source, the error may come from the language model’s parametric knowledge—what it learned in training. And sometimes a model makes up a detail with no basis at all; such hallucinations cannot be ruled out entirely.
Corrections are made where the wrong information appears: on the company’s own website, in its Organization markup, in its own profiles and business listings, and, where possible, on third-party sites. The goal is for the same up-to-date version to appear everywhere. Whether the corrections reach AI answers only becomes clear when the facts are checked again in later rounds of GEO monitoring. How quickly that happens depends, among other things, on whether an AI system retrieves content from the web when a question is asked or relies on what it learned in training.
A simple error log works well for your reports: which statement was wrong in which AI system, where it probably came from, where it was corrected, and whether it still appears in the next check. That way, each corrected statement can be documented on its own.