arXiv AI By Vin\'icius Gabriel Angelozzi, H\'eber H. Arcolezi

Where to Intervene? Benchmarking Fairness-Aware Learning on Differentially Private Synthetic Tabular Data

Read the original on arXiv AI →

arXiv:2607. 07471v1 Announce Type: cross Abstract: Machine learning models are increasingly deployed in high-stakes domains, raising concerns about both privacy and fairness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 4

Portable Causal Fairness Across Synthetic Data Generator Families

The paper demonstrates that causal fairness mechanisms can be applied across a wide range of synthetic data generators, including marginal‑based, GAN, and diffusion models, each with differentially private variants. By porting three fairness definitions to nine generators and testing them on Adult and COMPAS datasets, the authors show that the causal diffusion backbone consistently produces the fairest data releases while maintaining high fidelity. The fairness cuts have minimal impact on data quality, costing downstream classifiers only about $0.07$ to $0.15$ AUC on average, and adding privacy guarantees does not reduce fairness.

By Steven Golob, Sikha Pentyala, Martine De Cock
arXiv Machine Learning
1d ago

Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective

The paper argues that evaluating anonymity in synthetic data generation must focus on the generative model rather than just the resulting dataset. It interprets GDPR definitions of personal data and anonymization under realistic model-access scenarios, mapping these to state‑of‑the‑art privacy attacks. The authors conclude that synthetic data alone is insufficient for anonymization, and that Differential Privacy offers stronger protection than Similarity‑based Privacy Metrics.

By Georgi Ganev, Emiliano De Cristofaro