llms.txt is a proposal for a text file that tells large language models (LLMs) and AI agents which content on a website matters and where to find it. Jeremy Howard published the proposal in September 2024. The file usually sits directly under the domain, at /llms.txt, and is written in Markdown, which makes it easy to read for people and machines alike.

Structure

The proposal defines a fixed order:

  1. a level-one heading with the name of the website or project, the only required element
  2. a blockquote with a short summary
  3. optionally, further paragraphs or lists with background information
  4. sections under level-two headings that contain lists of links, each with a short note on what the linked page covers

By convention, a section titled “Optional” marks links that a system can skip when context is tight. The proposal also suggests offering important pages as clean Markdown versions. Some websites additionally provide an llms-full.txt that bundles the full text of their content in a single file. That is a convention, not part of the proposal.

What the file is for

Web pages contain a lot that is just noise for a language model: navigation, scripts, layout. An llms.txt provides a concise summary curated by the website itself instead. It fits easily into a model’s limited context window and helps AI agents find the right page. For GEO, it is one way to describe an offering clearly and in your own words.

Google checks llms.txt in Lighthouse

Google has added llms.txt to Lighthouse, the auditing tool that also powers PageSpeed Insights. It is part of the experimental Agentic Browsing category, which evaluates how well a website is built for AI agents and checks whether the file can be fetched and meets a few minimum requirements. Google describes the file as an emerging convention designed specifically for language models and AI agents: without it, agents may need more time to understand the structure and main content of a website. Since an llms.txt takes little effort to create, it is worthwhile for most websites.

What it does not do

llms.txt does not control access. Which crawlers may read a website is still defined in its robots.txt file. Nor does llms.txt replace structured data, which labels facts in machine-readable form directly on each page.

llms.txt is not a binding standard yet; each AI provider decides how its systems use the file. Like any single measure, it guarantees neither a mention nor an AI citation, but it is a simple way to describe an offering clearly for AI systems.