arXiv:2604. 23904v3 Announce Type: replace-cross Abstract: Synthetic tabular data are often evaluated by distributional similarity, privacy distance, or train-on-synthetic-test-on-real predictive performance, but these criteria do not ensure validity for causal inference.
By Yichen Xu
arXiv:2608. 02480v1 Announce Type: cross Abstract: With AI systems gaining more access to individuals' information, it is important to protect privacy when reporting statistical answers.
By Jinwon Sohn, Veronika Ro\v{c}kov\'a
With AI systems gaining more access to individuals' information, it is important to protect privacy when reporting statistical answers. Equally important is to privatize the reporting of uncertainty in such answers.
arXiv:2606. 08027v1 Announce Type: cross Abstract: Vertical federated learning (VFL) is a distributed learning paradigm that leverages vertically partitioned features across isolated parties without sharing raw samples; however, it remains vulnerable to active sample reconstruction attacks.
By Yongqi Jiang, Yansong Gao, Siguang Chen, Anmin Fu
arXiv:2606. 23741v1 Announce Type: cross Abstract: Causal reasoning, which encompasses the discovery of causal structures and the inference of causal effects, is fundamental to data-driven decision making.
By Xianjie Guo, Yuwei Wang, Guodu Xiang, Xiaoli Tang, Kui Yu, Han Yu, Qiang Yang
arXiv:2608. 15645v1 Announce Type: new Abstract: Transporting a causal conclusion from a source study population to a target one is a fundamental problem in causal inference.
By Yorgos Felekis, Paris Giampouras, Fabio Massimo Zennaro, Theodoros Damoulas
arXiv:2511. 14441v2 Announce Type: replace-cross Abstract: To distinguish Markov equivalent graphs in causal discovery, it is necessary to restrict the structural causal model.
By Daniel Klippert, Alexander Marx
arXiv:2602. 22083v2 Announce Type: replace-cross Abstract: Causal identification functionals often require integration over conditional densities of continuous variables, such as those arising in nonparametric identification theory of total and mediated causal effects in DAGs with hidden variables.
By Xiaxian Ou, Razieh Nabi
arXiv:2603. 10254v2 Announce Type: replace Abstract: Synthetic tabular data generation addresses data scarcity and privacy constraints in a variety of domains.
By Davide Tugnoli, Andrea De Lorenzo, Marco Virgolin, Giovanni Cin\`a
arXiv:2606. 16952v1 Announce Type: cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.
By Kareem Amin, Rudrajit Das, Alessandro Epasto, Adel Javanmard, Dennis Kraft, M\'onica Ribero, Sergei Vassilvitskii
We revisit the fairness notion of disparate impact for synthetic data generation (SDG), that assesses whether the utility of generated records is the same across sensitive groups. Our approach departs from existing work on fair SDG, that address the problem of correcting for undue biases in the observed distribution, hence redefining SDG as learning a distribution that is not that of the real data.
arXiv:2603. 12037v2 Announce Type: replace Abstract: Foundation models based on prior-data fitted networks (PFNs) have shown strong empirical performance in causal inference by framing the task as an in-context learning problem.
By Valentyn Melnychuk, Vahid Balazadeh, Stefan Feuerriegel, Rahul G. Krishnan