The rapid adoption of generative AI is pushing Copyright and AI Training to the centre of Europe’s legal debate, as policymakers, creators and technology companies confront a fundamental question: under what conditions can protected works be used to train artificial intelligence systems?
A new report from the European Audiovisual Observatory, Copyright and AI Training, examines that question from the perspective of European copyright law, with particular attention to the impact on creative sectors such as film, television and other audiovisual production.
Written by Diego de la Vega and published in July 2026, the report offers a clear overview of where the debate currently stands. Rather than providing a definitive answer, it looks at the rules already in place, the legal grey areas that remain and the competing interests of creators, rightsholders and AI developers.
The issue starts with the way AI systems are built. Training modern models requires huge volumes of data, and that material can include text, photographs, music, video and other works protected by copyright.
That makes the debate especially relevant for the audiovisual industry. As AI becomes more involved in writing, editing, dubbing, subtitling and image generation, questions around copyright are no longer limited to technology companies. They increasingly affect producers, creators and other professionals across the sector.
AI training puts copyright at the input stage
Much of the public debate around generative AI has focused on what an AI system produces. However, the Observatory concentrates primarily on what happens before the output is generated.
Training an AI model can involve collecting and processing large datasets, often through web crawling or scraping. The report identifies several stages in the process, from the initial ingestion and preprocessing of material to model training, generation, retrieval and the storage of certain information in logs or caches.
Different copyright rights may become relevant during these stages. These include the reproduction right, database rights, rights concerning computer programs and the neighbouring rights granted to press publishers.
The question becomes particularly important when protected works have been copied, extracted or processed without an individual licence from their rightsholders.
Text and data mining sits at the heart of the European debate
In Europe, much of the legal discussion revolves around the Copyright in the Digital Single Market Directive (CDSMD) and its exceptions for text and data mining (TDM).
Article 3 provides an exception for research organisations and cultural heritage institutions carrying out TDM for scientific research when they have lawful access to the material.
Article 4 is considerably broader. It allows reproductions and extractions of lawfully accessible works for TDM, but rightsholders can reserve their rights and therefore exclude their material from that exception.
This opt-out system has become one of the central points of contention surrounding AI training.
The Observatory notes that questions remain over whether the TDM provisions fully cover the much more complex technical process involved in training generative AI. Some legal scholars argue that TDM forms only one part of AI training and that the two concepts should not automatically be treated as equivalent.
There are also practical concerns about the effectiveness of opt-outs. Rights reservations for publicly available online material need to be expressed appropriately, often through machine-readable mechanisms.
Transparency is becoming increasingly important
The EU AI Act does not replace European copyright law. Instead, for copyright-related issues it works alongside the CDSMD.
One of its most significant contributions to the debate is transparency.
Providers of general-purpose AI models must maintain technical documentation covering areas including training methodologies and the type and provenance of data used. They must also make publicly available a sufficiently detailed summary of the content used to train their models.
The Observatory also highlights the General-Purpose AI Code of Practice, whose signatories include several major technology companies. Among other commitments, participating providers agree to respect technological restrictions such as paywalls and rights reservations communicated through mechanisms including robots.txt and other machine-readable protocols.
Together, these measures point towards a European system in which the ability to identify what material has been used — and whether rightsholders have reserved their rights — is becoming increasingly important.
Europe is not following the same path as the rest of the world
There is currently no single international approach to copyright and AI training.
The Observatory compares several systems. The EU relies on its TDM exceptions and opt-out mechanism, while the United States has no dedicated TDM exception and instead assesses AI training through the fair use doctrine.
The UK currently permits TDM through a statutory exception for non-commercial research and continues to debate possible reforms. Japan, meanwhile, has adopted a broader statutory exception covering commercial and non-commercial TDM in certain circumstances.
Canada and China follow different approaches again.
The result is a fragmented international environment for AI developers and rightsholders, even as the technology itself operates across borders.
Prompting creates another copyright layer
The report also examines what happens when users interact directly with AI systems.
Prompts can include instructions, images, texts or other material, potentially creating their own copyright questions. However, the legal position of prompts themselves remains unsettled.
The Observatory notes that current approaches generally distinguish between an idea or instruction contained in a prompt and the creative expression generated by the system. Questions therefore remain about how much human control is necessary before a user can claim authorship over AI-generated material.
Platform rules add another layer. Adobe, ChatGPT, Claude, Copilot and Midjourney all apply different terms governing how user inputs may be processed or used. Importantly, however, the report stresses that contractual terms between platforms and users do not override obligations imposed by European copyright law.
A framework still being tested
The central message of the Observatory’s study is that Europe already has a legal framework capable of addressing many aspects of AI training, but its application to generative AI is far from settled.
The CDSMD provides the copyright foundation, the AI Act adds transparency obligations, and rightsholders have mechanisms for reserving certain uses of their works. Yet uncertainty remains over whether existing TDM exceptions adequately cover every stage of AI training and whether the current opt-out system works effectively in practice.
Court decisions and the ongoing review of European copyright rules are therefore likely to play a major role in defining the next phase.
For film, television and the wider creative industries, the debate around Copyright and AI Training ultimately comes down to finding a workable balance: giving AI developers access to the data needed to build new technologies while preserving creators’ ability to control, license and potentially receive remuneration for the use of their work.
Continue reading: A24 and Google DeepMind Begin a New Experiment in AI and Filmmaking