On 27 September 2024, the Court in Hamburg held that LAION does not infringe the copyright of Robert Kneschke. Kneschke engages in stock photography. LAION is a not-for-profit organisation that makes databases available for the training of generative AI models.
In 2021, LAION created a data set with images on the basis of an existing American database. LAION downloaded the images from their storage locations (also known as scraping). Next, LAION checked with the help of software whether the downloaded image corresponded to the images in the American database. During this process, an illustration of Kneschke was also downloaded and checked. The ruling is contradictory in its description of the dataset. On the one hand, the dataset consists of a “table document” with “hyperlinks”, the Court said. On the other hand, the Court mentioned further down that the dataset consisted only of URLs (addresses of specific web pages or files on the internet) and a description of the images. The latter would mean than when someone wants to use this dataset to train an AI model, they will have to search for and download the images desired themselves.
The Court held that LAION did not infringe Kneschke’s copyright by downloading and checking the photograph. The composition of the dataset is included in the exception to copyright for “text and data mining with a view to scientific research” (Article 3 DSM Directive and § 60d Urheberrechtsgesetz). According to the Court, LAION composes the dataset for the purpose of scientific research. No immediate new knowledge or insights can be obtained by composing the dataset. Nevertheless, it is indeed a fundamental step towards the intended purpose of gaining knowledge at a later time. In coherence with the non-commercial objective it pursues as an organisation, LAION can rely on the exception and does therefore not infringe copyright.
The observations of the Court about the commercial text and data mining exception (Article 4 DSM Directive, § 44b Urheberrechtsgesetz) are also worth mentioning. With great restraint, the Court applies a broad interpretation of the term “machine-readable”. A “reservation formulated in natural language” is machine-understandable, according to the Court.
LAION can therefore compose the dataset. The Court did not go into the question as to whether the online dissemination of the dataset possibly copyright infringement.
It is doubtful whether the publication the dataset leads to copyright infringement. The Court remains unclear as to the composition of the dataset. Such facts may be relevant to the question whether publication of the dataset is a reproduction or a communication to the public.
If the dataset is assumed to consist merely of metadata, the dataset contains no easily downloadable collection of images. In this case, it is not obvious that the Court will rule that LAION is disclosing copyright-protected images. It is also not clear whether the making available of the dataset is a communication to a new audience. After all, the users themselves have to search for and download the images by means of the URL. If the dataset contains hyperlinks with which restrictive measures are circumvented, this may indeed constitute a communication to a new audience. This follows from the rulings of the European Court of Justice in the Svensson and GS Media cases.
Up to now, we do not have a clear answer to the questions above. A possible viewpoint in this assessment are the consequences for practice. This appears from recital 18 of the DSM Directive, inter alia. It is emphasized here that text and datamining is essential for, among other things, the development of new technologies by private entities. If a dataset for commercial exploitation is not accessible to the public, this may limit the use by others.
Future decisions will have to tell how a balance can be found between the protection of copyright holders on the one hand and the publishing of such datasets on the other hand.