arXiv:2606. 08179v1 Announce Type: cross Abstract: Subgraph counting is a fundamental problem in graph analysis.
By Xian Chen, Ruobing Bai, Pan Peng
arXiv:2602. 01607v3 Announce Type: replace-cross Abstract: Differentially private synthetic data enables the sharing and analysis of sensitive datasets while providing rigorous privacy guarantees for individual contributors.
By Rundong Ding, Yiyun He, Yizhe Zhu
arXiv:2602. 17284v2 Announce Type: replace Abstract: We consider the privacy amplification properties of a sampling scheme in which a user's data isused in $k$ steps chosen randomly and uniformly from a sequence (or set) of $t$ steps.
By Vitaly Feldman, Moshe Shenfeld
arXiv:2608. 11003v1 Announce Type: cross Abstract: In this work, we study the information bottleneck under perfect privacy, with particular emphasis on the active-rate regime, where the representation-rate constraint is binding and directly limits the achievable utility.
By Junle Zhong, Mohamad Assaad, Sreejith Sreekumar
arXiv:2604. 07486v3 Announce Type: replace-cross Abstract: Large language models (LLMs) have emerged as a powerful tool for synthetic data generation.
By Qian Ma, Sarah Rajtmajer
We revisit the fairness notion of disparate impact for synthetic data generation (SDG), that assesses whether the utility of generated records is the same across sensitive groups. Our approach departs from existing work on fair SDG, that address the problem of correcting for undue biases in the observed distribution, hence redefining SDG as learning a distribution that is not that of the real data.