Update of the legal playing field for personal data and AI with reference to the Digital Omnibus Package of the European Commission
AI models operatesa on the basis of data. This applies both to very large, often generatively deployable foundation models (such as those used for ChatGPT) and to smaller specialist (sector-specific) models designed for a specific purpose. These data often contain information about persons, and therefore fall under the regime of the General Data Protection Regulation (GDPR). For example, gigantic volume of data – often personal data – are 'scraped’ from the (public) internet to train foundation models. Earlier this year, tech parties Meta and LinkedIn announced that they would start using the personal data of the users of their social media platforms for the development and training of their own (foundation) models. It is striking that this processing does not rely on the basis of user consent. The legal basis that is relied on is the ‘legitimate interest’ under the GDPR.
Although the authoritative European Data Protection Board (“EDPB”) laid the foundation as early as 2024[1] for relying on the legitimate interest as a possible legal basis for the processing of personal data to train AI models, several data protection supervisors in the EU proved to be less unanimous on this point. For example, the Dutch DPA publicly warned[2] users to invoke their right to object to the proposed processing of their data for AI training by Meta and LinkedIn. And the Hamburg supervisory authority (Hamburgische Beauftragte für Datenschutz und Informationsfreiheit, “HmbBfDI”), after it initially labelling this activity as unlawful, ultimately decided not to pursue the enforcement proceedings against one of these big tech parties. This was partly in anticipation of an unequivocal EU framework on this issue. The recent bill of the European Commission to amend the GDPR, arising from the Digital Omnibus Package, looks like a serious step in the direction of this unequivocal EU framework.
The Digital Omnibus Package: overview and objectives
On 19 November 2025, the European Commission published the Digital Omnibus Package with the purpose of optimizing the application of the digital regulatory framework in the EU.[3] This proposal was drafted to make compliance with these regulations more cost-efficient, thereby achieving competitive advantages (including in the field of AI solutions). The starting point is that the level of protection of (EU) citizens and the intended objectives of each of these digital laws will not be affected. The proposal is mainly directed at EU organisations that are subject to this jumble of digital EU regulations, but also the big tech players who fall under the EU regulatory regime may profit from these benefits. Concretely, in order to achieve its goal, the European Commission has proposed a bill implementing a range of amendments to digital EU laws, including the GDPR.
Clarification of the term ‘personal data’ and pseudonymisation
Under the current version of the GDPR, personal data is defined as “any information relating to an identified or identifiable natural person”. Whether certain information can lead to the identification of a person depends on the means available to the party concerned or another person. These means cannot be in conflict with the law or require a disproportionate effort. This interpretation of the term is also called the ‘absolute’ doctrine.
In the Digital Omnibus, the European Commission follows a different path. In line with recent case law of the Court of Justice of the European Union[4], a context-dependent doctrine is chosen: the European Commission clarifies that information does not automatically have to be considered personal data by every party simply because another party is able to identify the person concerned. In the context of AI training, this means that a party that receives datasets containing pseudonymous personal data for this purpose, and does not have additional information about these persons, is not subject to the GDPR because these datasets cannot be considered personal data. The question of whether there is a legal basis for this processing, and to what extent this may be the legitimate interest, is then no longer relevant.
Expansion and clarification of the legal basis of ‘legitimate interest’ for AI training and use
The European Commission goes one step further in the Digital Omnibus: a new article will be included in the GDPR, which leaves no doubt that the processing of personal data for training and use of AI models and AI systems can be based on the legitimate interest. The controller must, meet all the conditions for invoking this basis.[5] Specifically, this means that a balancing of interests must be carried out to determine whether the controller's interest in training its AI model or using its AI system with personal data outweighs the interests of the data subjects.
In addition, strict security conditions must be observed, such as the application of data minimisation when selecting the sources from which the data are collected, but also during the training and testing of the AI system or model, and upholding an absolute, unconditional right of the data subject to object to the processing of their personal data in the context of AI.
It is clear that the European Commission has also taken the criticism of the Hamburg supervisor (partly) to heart, because it emphasizes that the above situation does not exist if it follows from EU or a Member State’s law that personal data can only be processed for AI on the legal basis of consent. In other words, legislators at Member State level have room to curb the legitimate interest as a legal basis for the processing of personal data for AI.
It is also worth mentioning that the European Commission has included a new, albeit limited, exception to the prohibition on processing for special categories of personal data (e.g. data concerning a person's health).[6] Specifically in the context of the development and use of an AI system or AI model, the processing of special categories of personal data is permitted on certain conditions. These conditions require that appropriate technical and organisational measures are taken to prevent the collection of special categories of personal data. If, nevertheless, special categories of personal data are discovered in the datasets used for training, testing or validating the AI system or model, and their removal would require a disproportionate effort, their processing is permitted. In that case, it is not necessary to obtain. It is a condition that the special categories of personal data cannot be used to create output and cannot be disclosed in any other way.
Practical consequences for organisations that process personal data for the training of AI models or the use of AI systems
The Digital Omnibus proposal provides clarity on the processing of personal data in an AI context. Under certain conditions, the training of AI models can be based on the legitimate interest under the GDPR. Obtaining consent, which is not always practical or feasible, may therefore be omitted. The same applies to the ‘remaining’ special categories of personal data that appear to have crept into datasets used.
A direct disadvantage of this method is that persons whose personal data ends up in an AI dataset through scraping will generally not be aware of this. Organisations are obliged to provide information on this, for example via the privacy statement on their website, but the likelihood that someone will independently or proactively consult this statement is limited.. The strong right to object granted to individuals by the Digital Omnibus proposal to stop this processing therefore merely exists on paper.
Finally, the question is whether, and if so, to what extent, the legislative proposal of the European Commission will remain intact. The European Parliament and the Council are currently deliberating on its content. Nevertheless, the current proposal shows an important shift in the approach of the European Commission in the field of digital regulation, which will affect all organisations involved in the training and use of AI with personal data.
[1] Guidelines 1/2024 on the processing of personal data based on legitimate interest. Although the public consultation of these Guidelines adopted on 08 October 2024 has ended, the final version has not yet been published by the EDPB. Their contents may therefore still be subject to change (https://www.edpb.europa.eu/our-work-tools/documents/public-consultations/2024/guidelines-12024-processing-personal-data-based_nl); Advice 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models (https://www.edpb.europa.eu/our-work-tools/our-documents/opinion-board-art-64/opinion-282024-certain-data-protection-aspects_nl).
[2] https://www.autoriteitpersoonsgegevens.nl/actueel/ap-kom-nu-in-actie-als-je-niet-wil-dat-meta-ai-traint-met-jouw-data; https://www.autoriteitpersoonsgegevens.nl/actueel/ap-bezorgd-over-ai-training-linkedin-en-roept-gebruikers-op-om-instellingen-aan-te-passen.
[3] https://digital-strategy.ec.europa.eu/en/library/digital-omnibus-regulation-proposal.
[4] ECJ 4 September 2025, case C-413/23 P, ECLI:EU:C:2025:645 (EDPS t. SRB).
[5] New article 88c GDPR.
[6] New Article 9 (2) (k) and 9 (5) GDPR.