Machine readability is about the difference between what people see on a website and what software can reliably take from it. A person sees at a glance that “from €49” below a product name is its price. Software recognizes this most reliably when it receives the information as text and the relationship is clear, for example from the page’s structure. AI systems read websites through such software: crawlers collect pages for training language models or as sources for answers, fetch tools load a page to answer a specific question, and AI agents operate websites in a browser.
What counts as machine-readable
EU law defines what a machine-readable format is. Under the EU Open Data Directive of June 2019, a file format is machine-readable if software can easily identify and extract specific data from it, down to individual statements of fact. Documents in a format from which data cannot be extracted, or only with difficulty, do not count as machine-readable. The directive covers data held by public sector bodies, but the standard carries over to a company website.
What makes a website machine-readable
Machine readability is not a single tool but a goal that several layers contribute to:
- Text: Key information such as services, prices, and contact details appears as text on the page, not only in images, videos, or graphics. Google lists making important content available in textual form among the fundamentals for its AI features, AI Overviews and AI Mode.
- Structure: Semantic HTML marks up headings, lists, tables, and page regions in the code as what they are, so software can recognize how a page is built.
- Meaning: Structured data states key facts such as a company’s name, address, or offering explicitly instead of leaving them to running text.
- Overview: An llms.txt file summarizes a website’s most important content for language models and AI agents. The llms.txt proposal also recommends offering key pages as Markdown versions, that is, as plain text without layout or scripts.
Before software can read a page, it has to be able to fetch it. Whether crawlers get access is a matter of crawlability; whether content only appears once JavaScript runs in the browser, and is therefore missing for some programs, is a matter of JavaScript rendering. Crawlability and machine readability both belong to On-Page GEO, the work on a brand’s own website.
Why machine readability matters for GEO
What software cannot extract from a page, an AI system can neither interpret correctly nor cite as a source. Microsoft states this explicitly in Bing’s Webmaster Guidelines, which also cover Copilot: AI answers based on web content, and their citations, depend on content that Bing can clearly interpret and verify.
Google states that its own AI features require no new machine-readable files, AI text files, special markup, or Markdown versions, because Google does not use them for these features. Many other programs that read websites can still benefit from machine-readable content, including other AI systems and AI agents. At the same time, Google’s auditing tool Lighthouse uses an experimental Agentic Browsing category to evaluate how well a website is built for machine interaction.
Machine readability often requires no new content, just existing information that is available as text, clearly structured, and labeled unambiguously in the code. It does not, however, automatically lead to a mention or an AI citation.
Machine readability in copyright law
Machine readability also plays a role in copyright law: under Section 44b of the German Copyright Act (UrhG), lawfully accessible works may be reproduced for text and data mining—their automated analysis, for example to compile AI training data—unless rights holders have reserved this use, and for works available online, such a reservation, known as a TDM opt-out, is only effective if it is made in machine-readable form. What counts as machine-readable here is legally disputed.