AI and Legislation | Legal Framework

AI and Copyright: Frameworks, Opportunities, and Pitfalls in the Netherlands and the EU

The rapid rise of generative AI systems has led not only to technological breakthroughs but also to fundamental legal questions. Content creators, tech companies, and end-users alike are struggling with the application of traditional copyright principles to artificial intelligence. Where does inspiration end and infringement begin? Who owns an AI-generated text or image? And how does European legislation relate to this new reality?

At its core, the debate on AI and copyright is divided into two major phases: the input phase (training the model) and the output phase (the content generated by the model). For companies and professionals working with AI, it is crucial to understand the distinction between these phases. In this article, we dissect the copyright framework for AI in the Netherlands and the European Union, so you know where you stand.

Rights to Training Data: The Input Phase

Large Language Models (LLMs) and image generators only function because they are trained on massive amounts of data. These datasets, such as the well-known Common Crawl, contain billions of web pages, books, articles, images, and forum posts. Virtually all of these materials are protected by copyright. The central question is: can an AI developer simply copy and analyze these protected works to train a model?

The Text and Data Mining (TDM) Exception

In the United States, tech companies often rely on the Fair Use doctrine, a flexible and open principle that in certain cases allows the use of copyrighted work without permission. In the European Union, and therefore also in the Netherlands, we do not have this system. The EU uses a closed system of specific statutory exceptions to copyright.

The most important exception for AI developers can be found in the European Directive on Copyright in the Digital Single Market (the DSM Directive from 2019). This directive introduced two specific exceptions for Text and Data Mining (TDM), which have been implemented in the Netherlands in the Dutch Copyright Act (Auteurswet, including Article 15o).

It is this Article 4 exception that currently forms the basis for much commercial AI development in Europe. In principle, it gives companies the right to scrape and use data for training, provided they have legal access to the source and no opt-out has been applied. This principle also strongly influences the development of open-source LLMs, where datasets sometimes operate in legal gray areas.

How Does the Opt-Out Work for Rightsholders?

As an author, publisher, or artist, you can prevent your work from being used for training AI models by making a so-called reservation of rights (opt-out). The law states that this reservation must be made in an "appropriate manner". For content that is publicly available online, the law explicitly requires that this reservation be made using machine-readable means.

In practice, this means that a simple legal disclaimer at the bottom of your website ("This work may not be used for AI training") may be insufficient for automated web crawlers. You must take technical measures. Commonly used methods include:

Managing these opt-outs is a challenge for many publishers, especially when data ends up in datasets via secondary sources or illegal copies, where the original opt-out is not visible.

Rights to the Output: The Generated Content

In addition to the question of whether the training was legal, the question arises about the output: who owns a text written by ChatGPT or an image generated by Midjourney? Can copyright rest on it?

The Requirement of 'Human Intellectual Creation'

The European Court of Justice ruled in the well-known Infopaq case that a work only enjoys copyright protection if it is the "author's own intellectual creation". This requires that the creator has made creative choices and that the work bears the stamp of their personality. This is the so-called originality requirement.

In both the Netherlands and the rest of the EU, there is a broad legal consensus that an "own intellectual creation" must, by definition, be human. Machines, algorithms, or animals cannot hold copyright. Therefore, the pure, raw output of an AI model generally falls into the public domain. It is not protected by copyright and can be freely copied or used by anyone.

Is Prompting Sufficient for Copyright?

A common argument is: "But I wrote a very complex prompt of up to five hundred words. Isn't the output my intellectual property then?"

Although this is a logical thought, current legal doctrine draws a sharp line here. Writing a prompt is often compared to giving a detailed instruction to an executing party (such as giving an idea to a painter). An idea, concept, or style is not protected by copyright in itself; only the concrete design (the final expression) is.

If the AI model itself determines the final design of the text or pixels based on probabilistic calculations, the human user lacks direct control over the exact, final expression. As a result, the required human creative input in the final result is missing. To claim copyright on AI content, a human author must significantly edit, select, rewrite, or arrange the generated material. The copyright protection then rests on the human edits, not on the AI-generated base. This theory, moreover, applies to all types of AI, including open weights models and closed commercial APIs.

Consequences for Businesses: Risks in Practice

The use of AI for content creation entails two major legal and commercial risks for companies: the risk of infringement and the loss of exclusivity.

1. The Risk of Copyright Infringement by the Output

Although the TDM exception may legitimize the *training* of the model, it does not cover infringements in the *output*. It is a known phenomenon that LLMs sometimes reproduce pieces of text exactly from their training data (this is called memorization or overfitting). If your employee enters a prompt and the AI happens to generate an exact copy of a protected article from De Volkskrant or The New York Times, and you publish this on your company blog, you may be infringing on the copyright of the original publisher.

In such a case, it is difficult to hide behind the AI model. As a publisher, you are responsible for what you make public. This emphasizes the absolute necessity of human quality control (human-in-the-loop) and plagiarism checks before AI content is deployed commercially.

2. No Exclusivity on Generated Materials

If you invest hundreds of thousands of euros in a marketing campaign that is fully generated using AI images without significant human editing, there is highly likely no copyright resting on these images. In concrete terms, this means that a direct competitor can download and use your campaign images for their own marketing, without you being able to take legal action against them via copyright. After all, the material belongs to the public domain.

Companies must therefore think strategically about which business-critical or brand-differentiating elements they have generated by AI, and where human (and thus protectable) design is essential. If you want to integrate AI into your business processes in a responsible manner, it is advisable to carefully study the general guide on AI implementation strategies.

Contractual Agreements with Suppliers and Freelancers

Given the legal uncertainties, companies must take proactive measures in their procurement and contract management. When working with advertising agencies, copywriters, software developers, or other external parties, it is of vital importance to contractually define the use of AI.

Without explicit agreements, you run the risk of paying for unprotected work, or being held liable for infringements committed by your supplier through careless AI use. Important elements to include in your General Terms and Conditions or framework agreements are:

The Impact of the AI Act on Copyright

In addition to existing copyright legislation, the new European AI Act (AI Regulation) also plays a crucial role. The AI Act itself does not create new copyrights, but it does introduce heavy transparency obligations for developers of so-called General Purpose AI (GPAI) models, such as OpenAI, Google, and Anthropic.

Under the AI Act, providers of these models must:

  1. Draw up and comply with a policy to respect European copyright law, in particular regarding the opt-out (TDM reservation) as described in Article 4 of the DSM Directive, regardless of where in the world the model is physically trained.
  2. Publish a sufficiently detailed summary of the training data used to train the model.

These requirements are intended to break the current asymmetry in information. Currently, authors simply do not know whether their work has been used in a model. The detailed summaries from the AI Act will enable rightsholders in the future to better check whether their rights have been violated and to take more targeted legal action.

Conclusion

The intersection between AI and copyright is dynamic and complex, but the foundations in the Netherlands and the EU are clearly anchored in legislation and case law. The Text and Data Mining exception offers AI developers room, provided they respect the (machine-readable) opt-outs of rightsholders. For the output, it is indisputable: without human, creative input, there is no question of copyright protection.

Companies embracing generative AI must be aware of the risks of infringement and the lack of exclusivity on purely machine-generated content. By implementing strict internal guidelines, clear contractual agreements with external partners, and careful human review, the benefits of AI can be safely and effectively utilized within the boundaries of the law.