Skip to content

GEO glossary

Key GEO terms, clearly explained.

The most important terms in GEO and AI visibility, each with a short definition and a clear explanation.

  • 73 terms

A

  • Accessibility tree

    The accessibility tree is a simplified version of a loaded web page that the browser derives from the page’s code, listing each relevant element with its role, name, and state. Screen readers and other assistive technologies rely on it, and AI agents that operate websites in a browser can read it as well.

  • Agent-friendly website

    An agent-friendly website is built so that AI agents can reach it, understand its content, and operate its controls, for example through semantic HTML, clearly labeled buttons and form fields, and a stable layout.

  • AI agent

    An AI agent is an AI system in which a language model pursues a goal over several steps and decides on its own which tools to use, such as a browser, the retrieval of web content, or application programming interfaces. This lets an AI agent, for example, compare offers or book an appointment on a person’s behalf.

  • AI answer opt-out

    An AI answer opt-out is a setting or machine-readable instruction that website owners use to indicate that AI systems should not use a website’s content as the basis for answers or show it as a source. With many providers, it is separate from the decision on whether the content may be used to train AI models.

  • AI citation

    An AI citation is a source reference in an answer generated by an AI system: a link, footnote, or source card that attributes information in the answer to a specific web page or document.

  • AI crawler

    An AI crawler is a program that fetches web pages on behalf of an AI provider, mainly to train AI models, to build an index from which AI systems draw sources for their answers, or to retrieve a page for a question a person is asking right now. Most major AI providers document a separate crawler name for each purpose, so websites can address each crawler individually in robots.txt.

  • AI index crawler

    An AI index crawler fetches and captures web pages in advance, independently of individual questions, so that an AI system can later draw on them, cite them as sources, and link to them in its answers. Examples include OpenAI’s OAI-SearchBot, Anthropic’s Claude-SearchBot, and Perplexity’s PerplexityBot.

  • AI referral traffic

    AI referral traffic consists of visits to a website that come from people clicking links in the answers of AI systems. Web analytics tools identify these visits by the referrer (the origin address the browser passes on) or by parameters in the link such as utm_source=chatgpt.com.

  • AI share of voice

    AI share of voice is a brand’s share of all mentions that it and a defined group of competitors receive in AI answers to a fixed set of prompts. A variant measures instead what share of the AI citations in those answers points to the brand’s website—within the competitor group or among all cited sources.

  • AI training opt-out

    An AI training opt-out is a machine-readable signal by which a website indicates that its content should not be used to train AI models, while other uses, such as appearing in AI answers, can remain allowed. It usually consists of robots.txt rules for the crawlers and tokens that AI providers document for training.

  • AI visibility

    AI visibility describes how often, in what context, and how accurately a brand or offering appears in the answers of AI systems such as ChatGPT, Gemini, Claude, or Perplexity.

  • Answer variability

    Answer variability is the tendency of AI systems to give different answers to the same question. The brands, facts, and sources an answer names can change each time the question is asked, with its phrasing, and over time.

  • Answer-first structure

    Answer-first structure is a writing pattern in which a page or section opens with the direct answer or key fact, followed by explanation, evidence, and detail, much like the inverted pyramid of news writing.

  • Authorship

    Authorship is the clear attribution of content to identifiable people or organizations, for example through bylines, author pages, and stated qualifications.

B

  • Bing AI Performance

    Bing AI Performance refers to the AI Performance report in Bing Webmaster Tools, which shows how often a website’s pages are cited as sources in Microsoft Copilot, AI-generated summaries in Bing, and select partner integrations. Microsoft introduced it as a public preview on February 10, 2026.

  • Bot management

    Bot management covers features of content delivery networks, firewalls, and hosting providers that detect automated traffic and allow, block, rate-limit, or charge for it. It works independently of robots.txt: a blocked crawler does not receive the page, even if robots.txt allows it.

  • Brand accuracy

    Brand accuracy describes how closely the statements an AI system makes about a brand match verified facts, for example about its offering, prices, locations, and contact details.

  • Brand fact consistency

    Brand fact consistency means that a brand’s core facts, such as its name, short description, offering, locations, contact details, and key figures, are stated the same way on its own website, on its own profiles, and in third-party sources.

  • Business listing

    A business listing, also called a directory listing, is a structured entry about a company in a directory or map service, or a company profile on a platform, with details such as its name, address, business hours, and category.

C

  • ChatGPT

    ChatGPT is OpenAI’s AI assistant. It answers questions from what its language models learned in training and, when useful or requested, can retrieve current content from the web and link to the sources it used.

  • ChatGPT search

    ChatGPT search is the ChatGPT capability through which OpenAI’s AI assistant retrieves current content from the web for an answer and can link to the sources it used; OpenAI introduced it under this name in October 2024. Depending on the question, ChatGPT decides on its own whether to use the web; the person asking can also start the retrieval deliberately.

  • Citability

    Citability describes how well a passage of web content can be retrieved, understood without its surrounding context, and attributed to its source, so that an AI system can use it in an answer and cite it.

  • Citation rate

    Citation rate is how often a website or an individual page is given as a source in AI answers, expressed either as the absolute number of citations in a period or as the share of answers that cite it among all answers to a defined set of questions (prompts).

  • Claude

    Claude is Anthropic’s AI assistant and the name of the language models behind it. Since 2025, Claude has been able to retrieve current content from the web when needed and cite the sources it used.

  • Claude-SearchBot

    Claude-SearchBot is the web crawler Anthropic says it uses to analyze online content to improve the relevance and accuracy of the answers Claude gives with its web search. Site owners can allow or block it with its own rule in robots.txt, independently of Anthropic’s other bots.

  • ClaudeBot

    ClaudeBot is the web crawler Anthropic uses to collect web content that could contribute to training its AI models. According to Anthropic, it honors robots.txt rules and supports the non-standard Crawl-delay field.

  • Content freshness

    Content freshness is how current a piece of content is and how clearly its publication and update dates are signaled to people and machines.

  • Crawlability

    Crawlability describes how well automated programs, such as the crawlers of AI providers, can discover and fetch the pages of a website. It depends on links, robots.txt rules, server responses, bot protection, and whether the content is in the delivered HTML or only appears after JavaScript runs.

E

  • Earned media

    Earned media is content about a brand that others publish without being paid for it, such as editorial coverage, independent reviews, and recommendations by experts or customers. The marketing term distinguishes it from paid advertising (paid media) and a brand’s own channels (owned media).

  • Entity

    An entity is a uniquely identifiable thing, such as an organization, a product, a person, or a place, that is described by its attributes and its relationships to other entities.

G

  • Gemini

    Gemini is Google’s family of AI models and the name of the AI assistant through which people use these models directly. The Gemini app answers questions from what its models learned and can draw on current content from Google’s web index and link to sources.

  • Generative AI performance report

    The generative AI performance report is a Google Search Console report that shows how often links to a website appeared in Google Search’s generative AI features, such as AI Overviews and AI Mode. It counts these appearances as impressions and breaks them down by page, country, device, and date.

  • Generative engine

    A generative engine is an AI system that answers a question by retrieving information from several sources and having a language model synthesize it into one response, usually with source references.

  • GEO (Generative Engine Optimization)

    GEO (Generative Engine Optimization) is the discipline of analyzing and improving how visible a brand or offering is in the answers of generative AI systems such as ChatGPT, Claude, or Perplexity.

  • GEO monitoring

    GEO monitoring is the continuous, repeated measurement of how often and how a brand or website appears in the answers of AI systems, compared with competitors and over time.

  • Google AI Mode

    Google AI Mode is a separate area of Google Search in which an AI system answers questions in a conversation and shows links to websites. It breaks a question into subtopics and researches them in parallel.

  • Google AI Overviews

    Google AI Overviews are AI-generated summaries that Google shows in Google Search for questions where its systems consider them helpful. They include links to web pages that support the information.

  • Google-Extended

    Google-Extended is a robots.txt product token with no crawler of its own. A rule for it determines whether Google may use content it crawls from a website to train future Gemini models and to ground answers, for example in Gemini Apps.

  • GPTBot

    GPTBot is OpenAI’s web crawler for content that may be used to train the company’s generative AI foundation models. When a site disallows GPTBot in its robots.txt, it indicates that its content should not be used for that training.

  • Grounding

    Grounding is anchoring a language model’s answer in external, usually current sources that the model receives when a question is asked, such as web pages, so that the answer’s statements can be traced back to those sources.

H

  • Hallucination

    A hallucination is a plausible-sounding statement, generated by a language model, that is factually wrong, not supported by its sources, or made up.

J

  • JavaScript rendering

    JavaScript rendering is the execution of a web page’s JavaScript by a browser or crawler to build the finished page. Content that only appears during this step is missing for programs that read only the HTML delivered by the server.

K

  • Knowledge cutoff

    The knowledge cutoff is the date up to which a language model’s training data reliably covers information. The model can only reliably take later events and changes into account if they are provided when it answers a question, for example because the AI system retrieves current content from the web.

L

  • Lighthouse Agentic Browsing

    Lighthouse Agentic Browsing is an experimental category in Google’s auditing tool Lighthouse that checks how well a web page is built for AI agents, for example through audits of the accessibility tree, layout stability, llms.txt, and WebMCP. Instead of a 0–100 score, it reports the share of checks passed.

  • LLM (large language model)

    A large language model (LLM) is an AI model that has been trained on vast amounts of text to predict how a text continues and that stores what it has learned in a very large number of numerical values called parameters. In most cases, it has also been trained to answer questions and follow instructions; models trained this way write the answers of AI systems such as ChatGPT, Gemini, Claude, and Perplexity.

  • llms.txt

    llms.txt is a Markdown file that a website provides for language models and AI agents: a compact, curated overview of its most important content, each item with a link and a short description.

M

  • Machine readability

    Machine readability means that content or data is structured so that software can reliably identify, interpret, and extract specific information without human help. For websites, it is a prerequisite for AI systems to capture their content correctly and use it in answers.

  • Mention

    A mention is the naming of a brand, product, or organization in a text without a source reference. The term covers both AI answers and content on other websites that AI systems learn from or retrieve for their answers.

  • Mention rate

    Mention rate is the percentage of AI answers to a defined set of questions (prompts) in which a brand is named at least once. It counts the brand’s appearance in the answer text, regardless of whether the answer gives a source.

  • Microsoft Copilot

    Microsoft Copilot is Microsoft’s AI assistant, available as a standalone app and built into Windows, Edge, and Microsoft 365 apps. When Microsoft Copilot draws on web content, it bases its answer on results from Bing and links to the sources it used.

O

  • OAI-SearchBot

    OAI-SearchBot is OpenAI’s web crawler for surfacing and linking to websites in ChatGPT search answers. A site indicates separately, through its robots.txt rule for GPTBot, whether its content should be used to train OpenAI’s models.

  • Off-Page GEO

    Off-Page GEO covers all measures outside a brand’s own website that influence whether and how AI systems mention and describe the brand or its offering in their answers. It includes building presence in independent media, on review platforms, in communities, in directories, and in reference works such as Wikipedia or Wikidata.

  • On-Page GEO

    On-Page GEO covers the measures on a brand’s own website or platform that make its content accessible, understandable, and citable for AI systems, from crawler access to semantic HTML, structured data, and clearly structured, self-contained text.

  • Organization markup

    Organization markup is schema.org structured data of the type Organization or one of its subtypes that describes a company in machine-readable form: its name, website, logo, address, contact details, and identifiers.

P

  • Parametric knowledge

    Parametric knowledge is the information a language model has stored in its parameters during training. The model can draw on it without looking anything up, unlike content an AI system retrieves from external sources at the moment a question is asked.

  • Perplexity

    Perplexity is an AI answer service that responds to questions in its own words and backs its answers with numbered source citations. By default, it retrieves the content for those answers when a question is asked, mainly from its own web index.

  • PerplexityBot

    PerplexityBot is Perplexity’s web crawler that collects websites so that Perplexity can show them in its answers and link to them. According to Perplexity, it is not used to collect content for training AI foundation models.

  • Prompt

    A prompt is the input a person or an application gives a language model to generate a response: a question, an instruction, or an entire conversation. In GEO, a brand’s visibility in AI answers is measured for specific prompts.

  • Prompt set

    A prompt set is the defined collection of prompts used to measure a brand’s visibility in AI answers repeatedly and comparably. The prompts are grouped, for example, by topic, buying stage, market and language, or persona.

  • Prompt tracking

    Prompt tracking is the automated, repeated submission of a defined set of prompts to AI systems, together with recording and analyzing their answers. It provides the data from which metrics on a brand’s visibility in AI answers are calculated.

Q

  • Query fan-out

    Query fan-out is a technique in which an AI system splits a question into several related subqueries, retrieves content for each of them, and combines the results into one answer.

R

  • RAG (retrieval-augmented generation)

    Retrieval-augmented generation (RAG) is a method in which an AI system retrieves documents or passages relevant to a question when the question is asked and has a language model generate its response from them. This allows the answer to include information beyond what the model learned in training.

  • robots.txt

    robots.txt is a text file at the root of a domain or subdomain that tells crawlers, by their user-agent token, which paths they may fetch. It is standardized as the Robots Exclusion Protocol in RFC 9309; its rules are a request to crawlers, not a form of access authorization.

S

  • Semantic HTML

    Semantic HTML is the use of HTML elements according to their defined meaning, such as headings, lists, tables, page regions, and buttons, so that software can recognize the structure of a page and the role of its content.

  • SSR (server-side rendering)

    SSR (server-side rendering) is the generation of a web page’s complete HTML on the server, so that its content can be read without running JavaScript. Static generation is a variant in which this HTML is produced in advance, when the website is built.

  • Structured data

    Structured data is machine-readable information in the code of a web page that describes its content with a shared vocabulary such as schema.org, for example a company, a product, or an article.

T

  • TDM opt-out

    A TDM opt-out, also called a TDM reservation, is a rights holder’s declaration that excludes their works from the general legal exception for text and data mining, the automated analysis of digital works. For works available online, German copyright law requires the reservation to be machine-readable; the underlying EU directive names machine-readable means for this.

  • Training crawler

    A training crawler is a web crawler that a provider runs to collect content for training AI models, such as OpenAI’s GPTBot, Anthropic’s ClaudeBot, or Mistral AI’s MistralAI-Training. Several providers give their training crawlers a name of their own, so websites can allow or block training in robots.txt separately from other purposes.

  • Training data

    Training data is the text and other content a language model learns from before it is put to use. It typically consists of filtered web content, datasets that are purchased, licensed, or openly available, and synthetic data generated by other AI models.

U

  • User-triggered fetcher

    A user-triggered fetcher is a program run by an AI provider that retrieves a specific web page because a person using an AI system is asking a question, sharing a link, or assigning a task—not as part of crawling the web on its own. Examples include ChatGPT-User, Claude-User, and Perplexity-User.

W

  • Web accessibility

    Web accessibility means designing and developing websites and other digital services so that people with disabilities can perceive, understand, navigate, and interact with them, including through assistive technologies such as screen readers.

  • Web index

    A web index is a database in which a provider stores the content of crawled web pages in processed form so that pages and passages matching a query can be found within moments. AI systems that draw on web content for their answers usually retrieve it from such an index.

X

  • XML sitemap

    An XML sitemap is a file in the sitemaps.org format that lists a website’s URLs, optionally with the date each page last changed. It helps crawlers discover pages and notice changes.