AI models, especially AI models for general purposes that can perform a wide range of different tasks (such as large language models), often use large training data sets for this purpose that also include personal data.
The General Data Protection Regulation (“GDPR”) prescribes that each data processing operation requires that a legally valid ground can be invoked. One of these grounds is the necessity for the defence of a legitimate interests pursued by the controller or by a third party (Article 6 (1) under f GDPR). The controller can only rely on this ground if the interest that is defended is “legitimate”, the processing of persona data is necessary, and the interests and fundamental rights of the data subjects (like the protection of their personal data) does not outweigh the interests of the controller.
We conclude that this ground is used ever more often – instead of consent – as a basis for the development, training and deployment of AI models. This is understandable, since for a valid reliance on consent would require obtaining the prior, specific, informed and unambiguous consent of each data subject of whom personal data are being processed for the AI model. In addition, this consent must be given freely, without intrinsic pressure or coercion. In practice this is not doable for most parties. The question whether, and on what conditions, relying on the legitimate interest is allowed in this context under the GDPR, was recently interpreted in more detail by the European Data Protection Board (“EDPB”) and several national data protection authorities.
EDPB guidelines and opinion: general review framework
In October 2024, the EDPB published the Guidelines 1/2024 on the processing of personal data based on legitimate interest.[1] The purpose of these guidelines is to offer controllers practical tools for testing and applying the legitimate interest as a ground for processing. For each of the three cumulative conditions that must be tested in a balancing of interests, concrete viewpoints and examples are given. For example, in line with case law of the Court of Justice of the European Union, it was established that a purely commercial interest to train the AI model may be justified.[2] Furthermore, for balancing the interests to the advantage of the controller it can be a factor that the controller has undertaken actions to make the processing known to the data subjects, wherefore it is assumed that the data subjects are reasonably aware of it. Even if the processing operation has a limited disadvantageous effect on the lives of the data subjects, this may be a factor to have the balancing exercise work out more to the advantage of the controller. The EDPB emphasized that this assessment must be documented and be performed prior to the processing. Therefore, in the development and/or training of an AI model it will usually be the provider who performs this balancing of interests, whether or not on instructions of the buyer (the deployer).
In December 2024, at the request of the Irish supervisory authority, the EDPB issued the Opinion 28/2024 for the processing of personal data in the context of AI models.[3] In this opinion the EDPB took the position that the use of personal data for AI training is not excluded in advance from the reliance on a legitimate interest, but that the legitimacy is highly dependent on the context, the nature of the data, the reasonable expectations of data subjects, and the deployment of measures to protect the fundamental rights of the data subjects. With this opinion the EDPB is trying to further European harmonisation by offering national supervisory authorities more clarity about the application of the legitimate interest in the context of AI models. In its recently updated manual for web scraping by private persons and private organisations[4] (a much-used phenomenon in development and training of AI models), the Dutch DPA (“AP”) has followed the viewpoints from the EDPB’s opinion in the weighing of interests for a valid reliance on the legitimate interest.
Remarkably, the EDPB has also thought about using AI models by third parties who wish to use those AI models for their own AI systems. Different scenarios are examined, in which the EDPB seems to give buyers of AI models (deployers) room to use these models lawfully nevertheless – on certain conditions – if the provider has processed personal data during development and training thereof without a valid ground under the GDPR.
Various approaches of supervisors within the EU
Several big tech businesses have already announced that they will start using personal data of their users to train AI models. They usually refer to the legitimate interest as a valid ground for the processing of personal data, such as the importance of improving products or developing new digital services. Among supervisory authorities within the European Union, opinions differ on the validity of this ground.
The Irish Data Protection Commission that acts as the lead supervisor for several of these big tech businesses has agreed to a reliance on legitimate interest as a ground for processing for the use of public user content for AI training, provided that extensive transparency measures, objection procedures and technical guarantees will be implemented.
The AP, on the other hand, has taken a more critical stance, but does not necessarily disapprove of a reliance on this ground. It has emphasized that the effectiveness of certain measures is only visible in practice, like building in a filter to strip data of personal characteristics and sensitive information, before they are used for training AI. The AP repeatedly called on consumers to use their right of objection in the context of AI trainings by big tech companies.
In Germany, there is division as well; while the data protection authority in Hamburg (Hamburgische Beauftragte für Datenschutz und Informationsfreiheit, “HmbBfDI”) initially objected to AI training on the basis of personal data on the basis of the legitimate interest, it eventually decided not to go ahead with the enforcement procedure against a big tech party, also in anticipation of a unilateral European framework on this point.
German judicial review and termination of national urgent proceedings
The HmbBfDI took into account a ruling of the Oberlandesgericht Köln before deciding to terminate the above-mentioned urgent enforcement proceedings. In a case brought by the Verbraucherzentrale Nordrhein-Westfalen, an interest group for consumers, the Court ruled on 23 May 2025 that the use of public (personal) data for AI training was not unlawful in the case at issue. It was decisive, among other things, that the controller had fulfilled information obligations, had offered objection procedures, and had taken measures to limit the impact on data subjects. The HmbBfDI indicated in a statement that an isolated national measure was not desirable, also in light of the pending European evaluation of these practices. For this reason, HmbBfDI does not want to be the only EU supervisor to issue a national preliminary ban on the AI training in that specific case. This reversal does not mean that any training of AI with personal data in Germany is fully shielded from risks.
EU supervisors do not yet see eye to eye
At the moment, there is no fully harmonised line (yet) within the European Union with regard to the reliance on the legitimate interest as a lawful ground for the processing of personal data in the context of AI. With its guidelines, the EDPB has admittedly provided a framework, but the application thereof remains dependant on the concrete circumstances of the processing and on the views of national supervisors. It is expected that in the coming period, further alignment will take place between the supervisory authorities, also on the basis of practical experiences. For example, the privacy interest group nyob filed a complaint in June 2024 to the AP, among other things, and Belgian and German supervisors against several big tech companies in connection with the processing of personal data for AI trainings and the DPC ceased proceedings against the platform X at the end of 2024, after X had promised to limit its use of personal data of its users for AI trainings.[5] The processing of personal data in the context of AI strongly stir up feelings (of supervisors).
Until there is one uniform EU framework, which may perhaps remain a Utopia, it remains important for organisations that develop, train and/or use AI models within their organisation to go about the
processing of personal data in that context carefully. A thorough, documented balancing will have to take place, in any case in line with the advice of the EDPB and the national (lead) supervisor, in order to assess whether the legitimate interest can be relied on as an appropriate ground for the GDPR.
Want to learn more about the processing of personal data in relation to AI models and systems? Or do you have other questions about the consequences of the AI Regulation taking effect? Please contact Laura Poolman or someone else from the Kennedy Van der Laan AI team.
[1] Although the public consultation of the Guidelines adopted on 08 October 2024 has already closed, the final version has not yet been published by the EDPB. Therefore, its contents may still be subject to change. https://www.edpb.europa.eu/our-work-tools/documents/public-consultations/2024/guidelines-12024-processing-personal-data-based_nl
[2] ECJ 4 October 2024, case C-621/22, Koninklijke Nederlandse Lawn Tennisbond (ECLI:EU:C:2024:857), ground 49.
[4] https://www.autoriteitpersoonsgegevens.nl/documenten/handreiking-scraping-door-particulieren-en-private-organisaties (updated on 2 April 2025).