AI training data summaries: what rights holders can learn and do next

Every provider of a general-purpose AI model placed on the EU market must publish a summary of the content used to train it, following the European Commission’s AI training data summary template. For publishers, music labels, photographers, software companies and any business with a valuable catalogue, that summary is the first public evidence of whether their works may have been used. This guide explains how to read each section of the template, what it will not tell you and what you can do next.

Key takeaways

  • Article 53(1)(d) of the EU AI Act requires the summary from all general-purpose AI model providers, including open-source ones.
  • Models placed on the market from 2 August 2025 need a summary at launch; older models have until 2 August 2027.
  • Providers must list the most relevant web domains they scraped: the top 10% by content size, or for SMEs the top 5% or 1,000 domains, whichever is lower.
  • The summary does not identify individual works. It is a starting point for a request, a licence negotiation or a claim, not proof in itself.
  • The AI Office’s enforcement powers over these rules apply from 2 August 2026, with fines of up to 3% of worldwide turnover or EUR 15 million.

What is the AI training data summary template?

The AI Act (Regulation (EU) 2024/1689) obliges providers of general-purpose AI models to “draw up and make publicly available a sufficiently detailed summary” of their training content, using a template from the AI Office (Article 53(1)(d)). The Commission published the explanatory notice and template on 24 July 2025; the version now on that page is Communication C(2025) 8311 of 5 December 2025.

According to the notice, the summary should cover all training stages, from pre-training to fine-tuning, and all data sources, protected or not. It must be published on the provider’s website and alongside the model in all its public distribution channels. If the provider keeps training the model, the summary should be updated every six months, or sooner if the new data requires a materially significant update.

How to read each section of the summary

Template section What the provider discloses What a rights holder can learn
1. General information Provider and EU representative, model versions, models it builds on, modalities, size within broad ranges, latest data collection date, languages Who to address, whether the model was trained before or after your content was published, and whether your language and media type are covered
2.1 Publicly available datasets Name and link of each “large” dataset (over 3% of the public data for a modality); general description of the rest Whether known datasets that may contain your works were used
2.2 Private third-party datasets Whether licensed datasets were used (with limited detail); other private datasets if publicly known Whether the provider licenses content, which signals that a licensing route exists
2.3 Crawled and scraped data Crawler names, purposes, behaviour (robots.txt, paywalls, captchas), collection period, type of sites, list of top domains Whether your domains appear, which crawler visited them and when
2.4 to 2.6 User, synthetic and other data Use of user interactions, models used to generate synthetic data, offline or human-labelled sources Whether content you supplied through the provider’s own services was reused
3.1 Text and data mining reservations Whether the provider signed the GPAI Code of Practice, and measures taken to honour opt-outs Which opt-out protocols the provider says it respects
3.2 Illegal content General description of filtering measures Whether infringing material was meant to be excluded

What the summary will not tell you

The notice is explicit about the limits. The template does not require disclosure of the specific works used, because Article 53(1)(d) asks only for a “summary” that is generally comprehensive rather than technically detailed. Licensed data gets lighter disclosure, since the rights holders are already parties to those deals. Content used only at inference time, for example through retrieval-augmented generation (RAG), is outside the mandatory sections unless the model actively learns from it.

The AI Office may check that the template has been filled in correctly, but it will not assess whether specific works were used. For domains not listed, the notice recommends that providers voluntarily tell rights holders, on request, whether content from a given domain was used. It also encourages mediation and points to the remedies in the IP Enforcement Directive.

What can rights holders do next?

  1. Find the summaries that matter. Start with the models your customers and competitors use and check the provider’s site and distribution channels.
  2. Search the domain list in section 2.3 for your websites, platforms where your content is hosted and known piracy sites carrying your catalogue.
  3. Cross-check crawler names and collection dates against your server logs and the date you introduced a text and data mining reservation.
  4. Compare section 3.1 with your opt-out method. Article 4(3) of the DSM Directive allows reservations “in an appropriate manner, such as machine-readable means” for content published online.
  5. Use the voluntary request route for domains not listed, and the provider’s complaints contact where it signed the Code of Practice, which commits signatories to a point of contact and a complaints mechanism for rights holders.
  6. Decide between licensing and enforcement. In court, Article 8 of the IP Enforcement Directive lets judges order information on the origin and distribution of infringing services.

If you need this done across catalogues and markets, our team for AI and digital asset protection can review the summaries, map the evidence and prepare the next step.

What this means for your business

For European rights holders, the summary turns a guessing game into a documented starting point. For Latin American and African catalogues, the key point is territorial: the obligation applies to models placed on the EU market, wherever the provider is established, so a Mexican publisher or a Nigerian music label can rely on summaries of models offered in Europe. Our recommendation is to keep a dated record of every summary you review, since providers must update them and earlier versions may disappear.

Where rights holders get this wrong

  • Expecting a list of works. The summary works at domain and dataset level; proof of use of a specific work needs further steps.
  • Checking only their own website, when much of their content sits on platforms, retailers or pirate sites.
  • Reserving rights with a human-language notice only, when machine-readable means are what crawlers read.
  • Waiting for older models: providers have until 2 August 2027 for those, but new models must publish at launch.
  • Moving to litigation without a licensing assessment, or the reverse. Our IP licensing and litigation team weighs both.

Frequently asked questions

Which AI providers must publish a training data summary?

All providers of general-purpose AI models placed on the EU market, including models released under free and open-source licences. Models placed on the market from 2 August 2025 need a summary at launch; for models placed before that date, providers must publish it by 2 August 2027 and justify any information gaps.

Will the summary tell me if my book or song was used?

Not directly. The template requires lists of large public datasets and top scraped domains, plus narrative descriptions, but not individual works. It can show that a domain or dataset holding your content was used, which supports a request to the provider, a licence negotiation or a court application for information.

When can the AI Office fine a provider for a missing or inadequate summary?

The notice states that the AI Office’s supervision and enforcement of the rules for general-purpose AI models starts on 2 August 2026. Under Article 101 of the AI Act, fines for providers of these models can reach 3% of annual worldwide turnover or EUR 15 million, whichever is higher.

Can IP Global Guard review AI training data summaries for our catalogue?

Yes. We review the summaries of relevant models, cross-check your domains, reservations and logs, and advise on licensing or enforcement. Where proceedings are needed in a given country, we coordinate qualified local counsel across Europe, Latin America and Africa from a single point of contact.

How IP Global Guard turns a training data summary into a strategy

IP Global Guard, the IP services line of META Channel Corporation Limited, works on copyright, AI and licensing in more than 25 jurisdictions across Europe, Latin America and Africa, with one strategy and one billing relationship. Within the same group, our regulatory colleagues advise on the EU AI Act.

Tell us which catalogues, websites and AI models concern you. We will review the published summaries and propose a route, from opt-out to licence or claim. Get in touch with our team.

This article is general information, not legal advice, and does not replace an assessment of your rights in a specific case.

Sources