The paper demonstrates that causal fairness mechanisms can be applied across a wide range of synthetic data generators, including marginal‑based, GAN, and diffusion models, each with differentially private variants. By porting three fairness definitions to nine generators and testing them on Adult and COMPAS datasets, the authors show that the causal diffusion backbone consistently produces the fairest data releases while maintaining high fidelity. The fairness cuts have minimal impact on data quality, costing downstream classifiers only about $0.07$ to $0.15$ AUC on average, and adding privacy guarantees does not reduce fairness.
By Steven Golob, Sikha Pentyala, Martine De Cock
arXiv:2607. 07471v1 Announce Type: cross Abstract: Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness.
By Vin\'icius Gabriel Angelozzi, H\'eber H. Arcolezi
Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness. Differential Privacy (DP) has become a gold standard for privacy-preserving data analysis, while fairness-aware mechanisms aim to mitigate discrimination against underrepresented groups.
We revisit the fairness notion of disparate impact for synthetic data generation (SDG), that assesses whether the utility of generated records is the same across sensitive groups. Our approach departs from existing work on fair SDG, that address the problem of correcting for undue biases in the observed distribution, hence redefining SDG as learning a distribution that is not that of the real data.
arXiv:2606.13105v2 Announce Type: replace
Abstract: We revisit the fairness notion of disparate impact for synthetic data generation (SDG), that assesses whether the utility of generated records is t...
By Paul Andrey, Micha\"el Perrot, Batiste Le Bars, Marc Tommasi
arXiv:2602. 05833v2 Announce Type: replace Abstract: There is a need for synthetic training and test datasets that replicate statistical distributions of original datasets without compromising their confidentiality.
By Laura Plein, Alexi Turcotte, Arina Hallemans, Andreas Zeller