Natural Identifiers for Privacy and Data Audits in Large Language Models
arXiv:2606. 24408v1 Announce Type: new Abstract: Assessing the privacy of large language models (LLMs) presents significant challenges.
arXiv:2606. 24408v1 Announce Type: new Abstract: Assessing the privacy of large language models (LLMs) presents significant challenges.
arXiv:2606. 16952v2 Announce Type: replace-cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.
The paper introduces ‘DP-SPIN’, a trusted‑curator framework that generates differentially private semantic plans for aggregate insight generation. ‘DP-SPIN’ maps each record to a bounded sparse nonnegative vector over pre‑defined semantic concepts, sums these vectors into a semantic sketch, and releases a noisy plan containing admitted concepts and their masses. The framework provides user‑level privacy by clipping each user’s contribution and ensures that the final summary is differentially private through post‑processing, with guarantees established under both add/drop and replacement adjacency.
arXiv:2606. 16952v1 Announce Type: cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.
arXiv:2512. 03238v2 Announce Type: replace-cross Abstract: High quality data is needed to unlock the full potential of AI for end users.
QuanText is a training‑free, large‑language‑model‑agnostic mechanism for releasing textual datasets that protects dataset‑level secrets such as the proportion of records with a particular diagnosis or gender. It perturbs both the secret distribution and correlated attribute distributions by selecting candidate release distributions close to the private empirical distribution and rewriting each text sample to match the chosen distribution using attribute‑related snippets. The method is inspired by the Statistic Maximal Leakage framework and, under idealized conditions, satisfies an SML guarantee, while empirical evaluations show a superior privacy‑utility trade‑off compared to existing data generation baselines.
The paper argues that evaluating anonymity in synthetic data generation must focus on the generative model rather than just the resulting dataset. It interprets GDPR definitions of personal data and anonymization under realistic model-access scenarios, mapping these to state‑of‑the‑art privacy attacks. The authors conclude that synthetic data alone is insufficient for anonymization, and that Differential Privacy offers stronger protection than Similarity‑based Privacy Metrics.
arXiv:2608.28934v1 Announce Type: new Abstract: Differential privacy (DP) has traditionally been used to provide theoretical upper bounds on an algorithm's stability to changing its training data. In...
arXiv:2606. 08259v1 Announce Type: new Abstract: This paper investigates the problem of generating synthetic tabular data with differential privacy (DP) guarantees, enabling data sharing in sensitive domains.
arXiv:2607. 07471v1 Announce Type: cross Abstract: Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness.
arXiv:2606. 10481v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning of large language models (LLMs) can exhibit problematic memorization of individual training examples.
arXiv:2310. 19043v3 Announce Type: replace-cross Abstract: Recent years have witnessed growing concerns about the privacy of sensitive data.