arXiv Machine Learning

Short paper: Models in the dark -- Rectification and erasure under GDPR in ML supply chains

arXiv:2606. 05946v1 Announce Type: new Abstract: The rights to rectification and erasure, as established under the General Data Protection Regulation (GDPR), are central to protecting individuals' privacy.

arXiv AI
Aug 18

Position: AI Governance Needs ISO-like Interoperability Protocols, Not Just Laws

The paper argues that AI governance should rely on ISO-like interoperability protocols rather than solely on jurisdiction-specific laws. It proposes standardized AI nutrition labels that include metrics for bias, energy usage, and data provenance to enable machine‑readable risk communication across borders. These protocols aim to reduce regulatory fragmentation, lower barriers for SMEs, and build public trust while allowing modular evolution with technology.

By Azmine Toushik Wasi, Mst Rafia Islam, Mahfuz Ahmed Anik, Taki Hasan Rafi, Md Manjurul Ahsan, Dong-Kyu Chae
arXiv Machine Learning
1d ago

Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective

The paper argues that evaluating anonymity in synthetic data generation must focus on the generative model rather than just the resulting dataset. It interprets GDPR definitions of personal data and anonymization under realistic model-access scenarios, mapping these to state‑of‑the‑art privacy attacks. The authors conclude that synthetic data alone is insufficient for anonymization, and that Differential Privacy offers stronger protection than Similarity‑based Privacy Metrics.

By Georgi Ganev, Emiliano De Cristofaro
arXiv AI
Jun 30

Towards Harnessing the Collaborative Power of Large and Small Models for Domain Tasks

arXiv:2504. 17421v2 Announce Type: replace-cross Abstract: Large language models (LMs) offer broad generalization capabilities but require vast amounts of data and computational resources for domain-specific tasks; small models (SMs), in contrast, are more efficient and tailored to specific domains yet lack general-purpose coverage.

By Yang Liu, Kejia Zhang, Bingjie Yan, Tianyuan Zou, Jianqing Zhang, Zixuan Gu, Xiangsen Chen, Jianbing Ding, Xidong Wang, Jingyi Li, Xiaozhou Ye, Ye Ouyang, Qiang Yang, Ya-Qin Zhang
arXiv Machine Learning
Jun 30

Efficient Unlearning with Privacy Guarantees

arXiv:2507. 04771v2 Announce Type: replace-cross Abstract: Privacy protection laws, such as the GDPR, grant individuals the right to request the forgetting of their personal data not only from databases but also from machine learning (ML) models trained on them.

By Josep Domingo-Ferrer, Najeeb Jebreel, David S\'anchez
arXiv AI
Jul 21

A Survey on Unlearnable Data

arXiv:2503. 23536v3 Announce Type: replace-cross Abstract: Unlearnable data (ULD) has emerged as an innovative defense technique to prevent machine learning models from learning meaningful patterns from specific data, thus protecting data privacy and security.

By Jiahao Li, Yiqiang Chen, Yunbing Xing, Yang Gu, Xiangyuan Lan
arXiv AI
Jul 8

Privilege and confidentiality in generative AI workflows

arXiv:2607. 05479v1 Announce Type: cross Abstract: Generative AI (GenAI) systems store and process client data in three distinct ways: in the model's parameters through training and memorisation, in the context window during a live session, and in knowledge databases for retrieval-augmented generation (RAG).

By V\'aclav Jane\v{c}ek, Thomas Melham