ITALY Trends and Developments Contributed by: Paolo Balboni, Luca Bolognini, Davide Baldini and Nicolò Maria Salvi, ICT Legal Consulting
“spiders” used by AI algorithm providers. The inquiry is prompted by the widespread practice of various AI platforms using web scraping to gather large amounts of information, including personal data, from websites managed by public and private entities for specific purposes, such as news and administrative transparency. The Garante invited trade associations, consumer groups, experts, and academic representatives to share their comments and contributions on security measures against the extensive collec - tion of personal data for algorithm training. Following the investigation, on 7 June 2024 the GPDP issued a guidance document on how to protect personal data published online from web scraping carried out by third parties for the pur - pose of training generative artificial intelligence models. The document provides data controllers who publish personal data online (eg, website publishers) with some recommended security measures to prevent or, at least, hinder web scraping. These measures are not mandatory: data controllers have a duty to autonomously assess, based on the principle of accountability, whether to implement them, taking into account elements such as the latest technology develop - ments and the costs of implementation. In the guidance document, the Garante suggests a number of concrete measures to be adopted, including: • the creation of reserved areas, accessible only upon registration, so as to remove data from public availability; • the inclusion of anti-scraping clauses in the terms of service of websites; • the monitoring of traffic to web pages, so as to identify any abnormal flows of incoming and outgoing data; and
• the implementation of specific measures against bots using, among others, the tech - nological solutions made available by the same companies responsible for web scrap - ing (eg, intervening on the robots.txt file). It is interesting to note that most of these meas - ures are in line with provisions stemming from other laws that regulate web scraping from other angles, such as the EU AI Act and the EU Copy - right Directive. For example, intervening in the robots.txt file in order to prevent or curtail web scraping from taking place is foreseen by the current draft of the Code of Practice for General- purpose AI (Article 56 AI Act), while the inclu - sion of anti-scraping clauses is aligned with the “text and data mining exception” pursuant to Article 4 of the Copyright Directive. This goes to show that the regulation of web scraping – as a fundamental pre-requisite for AI model training – brings an increasing degree of convergence between privacy and data protection, intellectual property, and AI regulation. In line with the above, in November 2024, the GPDP issued a warning against the Gedi Group, one of Italy’s biggest newspaper publishers, over its recent deal in relation to the sharing of Gedi’s newspaper archives with OpenAI in order to allow for the training of generative AI systems. In particular, the GPDP has reminded Gedi Group that, where the newspaper archives to be dis - closed to OpenAI contain personal data, such disclosure has to take place in compliance with GDPR, including as regards the right of each data subject to object to the disclosure. These initiatives further showcase the GPDP’s willingness and ability to accredit itself as a lead European regulator of AI systems, from a data protection perspective.
243 CHAMBERS.COM
Powered by FlippingBook