The paper argues that AI governance should rely on ISO-like interoperability protocols rather than solely on jurisdiction-specific laws. It proposes standardized AI nutrition labels that include metrics for bias, energy usage, and data provenance to enable machine‑readable risk communication across borders. These protocols aim to reduce regulatory fragmentation, lower barriers for SMEs, and build public trust while allowing modular evolution with technology.
By Azmine Toushik Wasi, Mst Rafia Islam, Mahfuz Ahmed Anik, Taki Hasan Rafi, Md Manjurul Ahsan, Dong-Kyu Chae
The paper argues that evaluating anonymity in synthetic data generation must focus on the generative model rather than just the resulting dataset. It interprets GDPR definitions of personal data and anonymization under realistic model-access scenarios, mapping these to state‑of‑the‑art privacy attacks. The authors conclude that synthetic data alone is insufficient for anonymization, and that Differential Privacy offers stronger protection than Similarity‑based Privacy Metrics.
By Georgi Ganev, Emiliano De Cristofaro
arXiv:2504. 17421v2 Announce Type: replace-cross Abstract: Large language models (LMs) offer broad generalization capabilities but require vast amounts of data and computational resources for domain-specific tasks; small models (SMs), in contrast, are more efficient and tailored to specific domains yet lack general-purpose coverage.
By Yang Liu, Kejia Zhang, Bingjie Yan, Tianyuan Zou, Jianqing Zhang, Zixuan Gu, Xiangsen Chen, Jianbing Ding, Xidong Wang, Jingyi Li, Xiaozhou Ye, Ye Ouyang, Qiang Yang, Ya-Qin Zhang
arXiv:2507. 04771v2 Announce Type: replace-cross Abstract: Privacy protection laws, such as the GDPR, grant individuals the right to request the forgetting of their personal data not only from databases but also from machine learning (ML) models trained on them.
By Josep Domingo-Ferrer, Najeeb Jebreel, David S\'anchez
arXiv:2503. 23536v3 Announce Type: replace-cross Abstract: Unlearnable data (ULD) has emerged as an innovative defense technique to prevent machine learning models from learning meaningful patterns from specific data, thus protecting data privacy and security.
By Jiahao Li, Yiqiang Chen, Yunbing Xing, Yang Gu, Xiangyuan Lan
arXiv:2607. 05479v1 Announce Type: cross Abstract: Generative AI (GenAI) systems store and process client data in three distinct ways: in the model's parameters through training and memorisation, in the context window during a live session, and in knowledge databases for retrieval-augmented generation (RAG).
By V\'aclav Jane\v{c}ek, Thomas Melham
arXiv:2609.01273v1 Announce Type: new
Abstract: Synthetic data has become a common component of machine learning research. While widely adopted, its use in privacy-sensitive contexts has quietly shif...
By Jiachen Zhao, Antonia Januszewicz, Taeho Jung
arXiv:2607. 10467v1 Announce Type: cross Abstract: Healthcare organizations often cannot freely centralize patient data because medical records are sensitive, regulated, and institutionally controlled.
By Sakshi Gorkhali, Jonesh Shrestha
arXiv:2502. 16184v3 Announce Type: replace Abstract: The EU Artificial Intelligence Act (AIA) establishes different legal principles for different types of AI systems.
By Henrik Nolte, Miriam Rateike, Mich\`ele Finck
arXiv:2606. 16952v1 Announce Type: cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.
By Kareem Amin, Rudrajit Das, Alessandro Epasto, Adel Javanmard, Dennis Kraft, M\'onica Ribero, Sergei Vassilvitskii
arXiv:2606. 16952v2 Announce Type: replace-cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.
By Kareem Amin, Rudrajit Das, Alessandro Epasto, Adel Javanmard, Dennis Kraft, M\'onica Ribero, Sergei Vassilvitskii
Machine unlearning has been extensively studied in response to growing privacy concerns and regulatory requirements. However, auditing whether unlearning algorithms have truly erased the influence of specific data remains an open challenge.