AI: input, prompt and output
(Generative) AI services like ChatGPT, Gemini, DALL-E and Midjourney are already widely known. These AI services consist of collected data (input) and when prompted, they can generate new content (output) in the form of a text or an image. Users can then decide to edit or publish the output. For users, this is a rather easy process. Moreover, no authors, photographers, models, etc. need to be engaged (and/or paid) to produce texts and images.
The output of AI is based on the input and the prompt of the user. The input – the collected data – is usually copyrighted, at least if it consists of works reflecting a degree of human creativity. Texts or images that are too commonplace and do not reflect creativity are not eligible for copyright protection.
Input and copyright
Copyright law grants authors the right to oppose the use – through publishing or reproducing (copying) – of their works. This right to oppose concerns not only direct copies, but also works of which the overall impression corresponds to the work of the author.
The input of an AI service consists of data – works – of authors. If this data is copyrighted and has been reproduced without the author’s consent, this may constitute copyright infringement. However, this cannot be concluded in general, since it is a factual matter that has to be assessed for each individual AI service. In addition, copyright may be set aside by the limitations provided in the Dutch Copyright Act.
Copyright is not absolute. No reliance can be made on copyright for (among other things) ‘temporary acts of reproduction’ and for text and data mining (known as the ‘TDM exception’), provided that these not contrary to a normal exploitation of works, and the legitimate interests of the author are not unreasonably prejudiced.
Limitations to copyrights: temporary reproduction and text and data mining
Temporary reproduction
Input as used by an AI service may be a temporary reproduction, which would prevent authors from making a successful reliance on copyright.Temporary reproductions are reproductions that:
- are transient or incidental;
- are an integral and essential part of a technological process;
- have as their sole purpose a lawful use of a work; and
- (iv) have no independent economic significance.
The question is whether the processing of the input of an AI service is a temporary reproduction within the meaning of the Copyright Act. For example, AI can be used to rewrite a complete article. This article can subsequently be published in its entirety (without the author's consent). In that case, it will have to be assessed whether this affects the normal exploitation of the work and the interests of the author who wrote the original article. As this is not a temporary reproduction, the above limitation will not apply.
Text and data mining
The input may also be covered by the limitation of text and data mining, which prevents the author from relying successfully on copyright. Text and data mining is an automated analytical technique aimed at analysing text and data – data of authors – in digital form, in order to generate information. AI is a form of text and data mining.
In order for a reliance on this limitation to be successful, the input of AI must have been obtained via lawful access. If the input comes from data behind a paywall, this limitation may therefore not apply. Texts and images containing a notice that the copyright is reserved to the author also do not fall under this limitation of copyright law. Such a reservation must be in ‘machine-readable form’. It is not clear yet what is meant by this; it is possible that an addition to the text or the use of a robot.txt file suffices (a file on website that prevents ‘crawling’ in practice, such as exclusion by Google). Websites that offer content, like news websites, often post a copyright reservation, which may cause the text and data mining limitation not to apply. Moreover, this limitation also cannot be contrary to a normal exploitation of the work and cannot unreasonably prejudice the legitimate interests of the rightholder.
Who to address?
It is not an easy task to prove that an AI service has used certain data. The input of AI usually consists of data from third parties, but that still does not prove that AI actually uses or has used your data.
Comparing the published output to the data of an author could be a way for proving this. If the output corresponds too much to that data, it could be argued that there is copyright infringement. However, it is usually assumed that little of the input can be recognized in the output. This leads us to the practical question of who should be addressed? Since the user of AI may have entered the full text or image in the prompt, it is doubtful whether the maker of the AI service can successfully be held liable.
It can therefore not be stated with certainty whether the processing of input in AI is an act of infringement according the Dutch Copyright Act. This question has not yet come up before a Dutch court, and it may also differ per AI service. It is rumoured that AI services are currently negotiating license deals with content makers (such as Reddit, Wordpress and Tumblr) in order to prevent legal conflicts.
Recent AI cases
In the US, several lawsuits are pending on the use of input against AI services, such as OpenAI and Microsoft. You will find an overview of these cases here (in Dutch). We will have to await their outcomes.
Meanwhile, a number of cases are also pending that deal with the question whether the output of AI is copyrighted. Recently, the US Copyright Office (USCO) registered a copyright for an AI-generated compilation of texts in a book written by a user of AI. The USCO had originally rejected the registration, whereupon the AI user lodged an appeal. On appeal, the USCO held that a copyright could indeed be registered for the book. This registration concerned the “selection, coordination, and arrangement of text generated by artificial intelligence”. This means that nobody may copy the book without consent, but that the sentences and paragraphs as such are not protected by copyright.
In Europe, on 11 October 2023 a Czech court ruled on the question whether an image generated by the AI service DALL-E was copyright-protected. The image had been generated by a short prompt that is freely translated as: “create a visual representation of two parties signing a business contract in a formal setting, for example in a commercial room or in the office of a law firm in Prague. Show only the hands.” The AI user who had entered the prompt claimed that his copyright had been infringed, because someone else was using this image. The Czech court dismissed the claim because:
- there was no evidence that the claimant was the actual author of the image with the prompt in question; and
- an image generated by AI cannot be copyrighted, because the copyright is vested in an individual, and the image was created by AI.
This ruling is regarded as the first ruling by a European court in the field of copyright and AI.
We will keep a close eye on developments in the AI field and will report on them on Mediareport.