
The European Union’s AI Act reached full enforcement on 2 August, bringing into force a requirement that developers of general-purpose artificial intelligence models publish a detailed summary of the data used to train them.
The obligation is the first binding transparency mandate of its kind, and gives copyright holders a formal route to see what material a model was built on. Models placed on the market before the deadline have until 2 August 2027 to comply.
What must developers disclose?
Article 53 of the regulation requires a sufficiently detailed summary of training content. The European Commission’s AI Office published a template setting out what that means in practice.
Providers must state the data modalities involved and the size of the dataset, identify the major public datasets used, and give narrative descriptions of licensed material, scraped material, user-generated content and synthetic data.
| Stage | Date |
|---|---|
| Regulation in force | 2 August 2025 |
| Full enforcement | 2 August 2026 |
| Deadline for pre-existing models | 2 August 2027 |
What does the EU AI Act require on training data?
Under Article 53 of the European Union’s AI Act, Regulation 2024/1689, providers of general-purpose artificial intelligence models must publish a sufficiently detailed summary of the content used to train them. Obligations for general-purpose models took effect on 2 August 2025 and reached full enforcement on 2 August 2026, with models already on the market before that date required to comply by 2 August 2027. The European Commission’s AI Office published a template in July 2025 specifying that disclosures must cover data modalities, dataset size, identification of major public datasets, and narrative descriptions of licensed, scraped, user-generated and synthetic data sources. It is the first binding transparency requirement globally to give copyright holders enforceable visibility into the datasets used to train AI models.
Why is this contested?
The requirement lands amid unresolved litigation over whether training on copyrighted work is lawful at all. Developers have argued that disclosure exposes commercially sensitive methods and invites claims; rights holders argue they cannot enforce copyright over material they cannot see.
Courts in the United States have so far split the question. A federal court ruled that training on copyrighted books can constitute fair use but that storing pirated copies does not. That case, brought by authors against Anthropic, settled for $1.5bn, among the largest copyright settlements recorded in the United States.
What else is in the courts?
The New York Times is suing OpenAI and Microsoft over the use of its articles. Universal Music Publishing Group, Concord and ABKCO filed a $3.1bn claim against Anthropic in January. Disney and other studios are pursuing the Chinese developer MiniMax.
Reuters has reported that rulings due this year could either clarify how fair use applies to AI training or deepen the uncertainty. The first jury trial in an AI copyright case is scheduled for September.
Some rights holders have settled instead. Disney agreed in December to invest $1bn in OpenAI and license its characters for the Sora video generator, and Warner Music resolved claims against the music generators Suno and Udio.

Leave a Reply