We revisit the fairness notion of disparate impact for synthetic data generation (SDG), that assesses whether the utility of generated records is the same across sensitive groups. Our approach departs from existing work on fair SDG, that address the problem of correcting for undue biases in the observed distribution, hence redefining SDG as learning a distribution that is not that of the real data.
Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness. Differential Privacy (DP) has become a gold standard for privacy-preserving data analysis, while fairness-aware mechanisms aim to mitigate discrimination against underrepresented groups.
arXiv:2607. 07471v1 Announce Type: cross Abstract: Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness.
By Vin\'icius Gabriel Angelozzi, H\'eber H. Arcolezi
The paper investigates how to maintain causal fairness when releasing synthetic data by applying the DECAF framework across nine different synthetic data generators from three families (marginals-based, GAN, and diffusion) and three levels of differential privacy. Experiments on Adult and COMPAS datasets show that the causal diffusion backbone consistently produces the fairest data with fidelity comparable to the marginals tier, while the fairness cuts have minimal impact on downstream classifier performance and do not degrade privacy guarantees.
The paper demonstrates that causal fairness mechanisms can be applied across a wide range of synthetic data generators, including marginal‑based, GAN, and diffusion models, each with differentially private variants. By porting three fairness definitions to nine generators and testing them on Adult and COMPAS datasets, the authors show that the causal diffusion backbone consistently produces the fairest data releases while maintaining high fidelity. The fairness cuts have minimal impact on data quality, costing downstream classifiers only about $0.07$ to $0.15$ AUC on average, and adding privacy guarantees does not reduce fairness.
By Steven Golob, Sikha Pentyala, Martine De Cock
arXiv:2602. 05833v2 Announce Type: replace Abstract: There is a need for synthetic training and test datasets that replicate statistical distributions of original datasets without compromising their confidentiality.
By Laura Plein, Alexi Turcotte, Arina Hallemans, Andreas Zeller
arXiv:2507. 19700v2 Announce Type: replace Abstract: We propose a new framework for generating tabular synthetic datasets via disjoint generative models.
By Anton Danholt Lautrup, Muhammad Rajabinasab, Tobias Hyrup, Arthur Zimek, Peter Schneider-Kamp
arXiv:2606. 09865v1 Announce Type: new Abstract: Privacy and data sharing are often in tension.
By Manel Slokom, Malek Slokom, Thierno Kante
arXiv:2604. 07486v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have emerged as a powerful tool for synthetic data generation.
By Qian Ma, Sarah Rajtmajer
arXiv:2606. 16952v2 Announce Type: replace-cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.
By Kareem Amin, Rudrajit Das, Alessandro Epasto, Adel Javanmard, Dennis Kraft, M\'onica Ribero, Sergei Vassilvitskii
arXiv:2606. 20461v1 Announce Type: new Abstract: Machine learning models have been shown to exhibit discriminatory outcomes or degraded performance for individuals at the intersection of multiple sensitive attributes, such as race and gender.
By Bruno Scarone, Alfredo Viola, Ren\'ee J. Miller
arXiv:2605. 11170v2 Announce Type: replace Abstract: Noise-based certified machine unlearning currently faces a hard ceiling: the noise magnitude required to certify unlearning typically destroys model utility, particularly for large-scale deletion requests.
By Ahmed Mehdi Inane, Vincent Quirion, Gintare Karolina Dziugaite, Ioannis Mitliagkas