The TDM opt-out, where TDM stands for text and data mining, is a copyright tool that matters for every website whose content AI providers can collect. The law generally allows lawfully accessible works, such as freely available text and images, to be reproduced for TDM, the automated analysis of works to extract information, for example about patterns, trends, and correlations. According to the EU AI Act, such techniques may be used extensively to train large AI models. A TDM opt-out is the legal side of the decision to keep content out of the training data of future AI models; the practical routes are what an AI training opt-out is about.

What the law says

The basis is the EU Directive on copyright in the Digital Single Market (DSM Directive) of April 2019. Its Article 4 allows TDM of lawfully accessible works only as long as rights holders have not expressly reserved this use in an appropriate manner, for example by machine-readable means for content made publicly available online. The directive’s recitals, the explanatory statements that precede its articles, add that only machine-readable means should count as appropriate for content made publicly available online, and that a reservation should not affect uses other than TDM. Germany implemented this rule in Section 44b of its Copyright Act (UrhG): for works accessible online, a reservation is only effective if it is made in machine-readable form.

The EU AI Act builds on this. Providers of general-purpose AI models, which typically include large language models, must have a policy to identify and comply with such reservations, including through state-of-the-art technologies. According to the regulation’s recitals, where rights holders have expressly reserved their works in an appropriate manner, providers need their authorization to carry out TDM on them.

Whether the TDM exception covers the training of large language models at all is a question for the Court of Justice of the European Union (CJEU): in 2025, a Hungarian court referred it, among other questions, in a dispute between a press publisher and Google (case C-250/25). As of early October 2026, no judgment had been delivered.

What counts as machine-readable

The law does not specify what machine readability requires for a reservation. The recitals of the DSM Directive give metadata and the terms and conditions of a website or service as examples.

How to interpret this is disputed, as a case between a photographer and the association LAION shows. To build a dataset that can be used to train AI models, LAION had downloaded one of the photographer’s images, among others, from a photo agency’s website. The agency had prohibited access by bots and other automated programs only in natural language, in its terms of use and in its website’s source code. In September 2024, the Regional Court of Hamburg expressed the view that such a reservation could suffice because AI applications can understand natural language, but it based its ruling on a different provision. In December 2025, the Higher Regional Court of Hamburg instead found that the photographer had not shown the reservation to be machine-readable at the relevant time of use in 2021. The Federal Court of Justice heard the case in September 2026 (case I ZR 281/25), indicated, according to reports on the hearing, that it might refer questions of interpretation to the CJEU, and has scheduled its decision for December 17, 2026.

A statement in a website’s legal notice or terms of use makes a reservation understandable to people. As of early October 2026, neither the Federal Court of Justice nor the CJEU had ruled on whether such text on its own counts as a machine-readable reservation.

Reservations in robots.txt

The EU’s Code of Practice for general-purpose AI models names one format explicitly: robots.txt. In this voluntary code, published in July 2025, signatories commit to using crawlers that read and follow robots.txt as specified in the RFC 9309 standard when they collect data for TDM and for training their models. Crawlers are programs that fetch web pages automatically. Signatories include OpenAI, Google, and Anthropic. For example, the lines “User-agent: GPTBot” and “Disallow: /” tell OpenAI’s training crawler, GPTBot, not to fetch any page of the site. Signatories that also operate an online search engine as defined in the EU’s Digital Services Act are further encouraged to ensure that websites declaring a reservation are not, as a direct result, put at a disadvantage in the provider’s index.

A rule in robots.txt addresses either a token—the name a provider defines for a crawler or for a specific use—or, with “User-agent: *”, every crawler that has no rules of its own in the file. It can therefore target a purpose such as AI training only through the tokens a provider names for that purpose: to keep content out of AI training, a site addresses each provider’s training crawlers and the tokens it documents specifically for training.

There are also formats designed specifically for reservations. TDMRep, for example, is a protocol that an open community group at the web standards body W3C first published in 2022 (current version: May 2024) and that is not an official W3C standard; among other places, it declares the reservation in a dedicated file on the server, in the server’s response, or in a page’s HTML code. Cloudflare, a provider of network services for websites, declares restrictions that sites state with its Content Signals, an additional line in robots.txt, to be express reservations under Article 4 of the DSM Directive; the crawler documentation of OpenAI, Google, Anthropic, and Perplexity does not say whether they honor these signals (as of October 2026). The Code of Practice provides for honoring such additional protocols if standardization organizations have adopted them, or if they are state of the art, widely adopted by rights holders, and generally agreed at EU level.

TDM opt-outs and AI answers

A TDM opt-out is a reservation against TDM. Whether a website can appear as a source in AI answers depends mainly on which crawlers its robots.txt allows. OpenAI, Anthropic, and Perplexity collect sources for answers with their own AI index crawlers, such as OAI-SearchBot, which have rules of their own; Perplexity says its crawlers do not collect content for training AI foundation models. With OpenAI and Anthropic, a site can therefore handle training and answers separately: blocking the training crawlers does not shut out the AI index crawlers. Google is different: the Google-Extended token covers both the training of future Gemini models and grounding in the Gemini Apps, that is, using content as the basis for answers.

If a site also blocks the crawlers that collect sources for answers, it largely forgoes having those AI systems use its pages as sources for answers and show them as AI citations. That is an AI answer opt-out: a decision that goes beyond a TDM opt-out and costs visibility in the answers it applies to. Even then, robots.txt does not cover every retrieval: not every provider applies it to a User-triggered fetcher, which retrieves pages because a person asks for them in a chat.

This page explains the legal situation in general terms and is not legal advice.