Disjoint Generation of Synthetic Data
arXiv:2507. 19700v2 Announce Type: replace Abstract: We propose a new framework for generating tabular synthetic datasets via disjoint generative models.
arXiv:2606. 09865v1 Announce Type: new Abstract: Privacy and data sharing are often in tension.
arXiv:2507. 19700v2 Announce Type: replace Abstract: We propose a new framework for generating tabular synthetic datasets via disjoint generative models.
arXiv:2606. 08372v1 Announce Type: cross Abstract: Synthetic data is increasingly promoted as a privacy-preserving substitute for releasing sensitive tabular records, yet its central adversarial threat ("reconstruction", the recovery of an individual's hidden attribute values from a synthetic release and a handful of known quasi-identifiers) has been studied only in scattered, hard-to-compare settings.
arXiv:2608.29674v1 Announce Type: new Abstract: Sharing tabular data in high-stakes domains is constrained by privacy regulations. Synthetic data offer a promising alternative, but deep generative mo...
arXiv:2603. 10937v2 Announce Type: replace Abstract: The use of synthetic data has become increasingly popular as a privacy-preserving alternative to sharing real datasets, especially in sensitive domains such as healthcare, finance, and demography.
arXiv:2606. 16952v2 Announce Type: replace-cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.
The paper argues that evaluating anonymity in synthetic data generation must focus on the generative model rather than just the resulting dataset. It interprets GDPR definitions of personal data and anonymization under realistic model-access scenarios, mapping these to state‑of‑the‑art privacy attacks. The authors conclude that synthetic data alone is insufficient for anonymization, and that Differential Privacy offers stronger protection than Similarity‑based Privacy Metrics.
arXiv:2602. 05833v2 Announce Type: replace Abstract: There is a need for synthetic training and test datasets that replicate statistical distributions of original datasets without compromising their confidentiality.
arXiv:2606. 16952v1 Announce Type: cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.
arXiv:2606.26403v2 Announce Type: replace Abstract: Foundation-model research increasingly needs data about people: user state, personal histories, relationships, contact-like fields, documents, and...
arXiv:2606. 10481v1 Announce Type: cross Abstract: Parameter-efficient fine-tuning of large language models (LLMs) can exhibit problematic memorization of individual training examples.
arXiv:2607. 07471v1 Announce Type: cross Abstract: Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness.
arXiv:2512. 03238v2 Announce Type: replace-cross Abstract: High quality data is needed to unlock the full potential of AI for end users.