TabCausal is a causal discovery foundation model that learns to map datasets directly to causal graphs by pretraining across diverse causal environments. It uses a dynamic task construction strategy to expose the model to varied graph priors, mechanisms, noise models, dimensions, sample sizes, and intervention regimes, improving transferability from observational and mixed‑interventional data. On large synthetic benchmarks and a new protocol‑guided semantic benchmark, TabCausal outperforms many classical baselines and shows robust structure recovery, especially when interventional evidence is available.
By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
arXiv:2603. 10254v2 Announce Type: replace Abstract: Synthetic tabular data generation addresses data scarcity and privacy constraints in a variety of domains.
By Davide Tugnoli, Andrea De Lorenzo, Marco Virgolin, Giovanni Cin\`a
The paper investigates how synthetic pretraining priors used in tabular foundation models (TFMs) influence downstream performance. By reconstructing the synthetic data generators of four TFMs and comparing their generated tasks to two popular tabular benchmarks using structural descriptors, the authors measure structural coverage and normalized density. They find that some generators provide broader and denser support for benchmark tasks, and that stronger synthetic-to-benchmark support generally correlates with better model performance.
By He Zhao, Ryan Thompson, Daniel M. Steinberg, Ashfaqur Rahman, Edwin V. Bonilla, Cheng Soon Ong
CausalArena is a new benchmark designed to evaluate causal discovery methods in the era of foundation models. It unifies synthetic structural causal models (SCMs), semantically grounded SCMs, and formula‑grounded SCMs, while also including real‑world datasets for external validation. Experiments show that performance rankings vary widely across different SCM families and protocols, indicating that strong results on one benchmark do not necessarily transfer to others.
By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
arXiv:2603. 10823v2 Announce Type: replace-cross Abstract: Deep generative models can help with data scarcity and privacy by producing synthetic training data, but they struggle in low-data, imbalanced tabular settings to fully learn the complex data distribution.
By Xiaofeng Lin, Seungbae Kim, Zhuoya Li, Zachary DeSoto, Charles Fleming, Guang Cheng
CausalArena is a unified, evolvable benchmark designed to evaluate causal discovery methods across diverse structural causal models (SCMs). It incorporates synthetic SCMs for controlled structural variation, semantic operational SCMs for human-auditable environments, and formula-grounded SCMs to test discovery under explicit scientific mechanisms, along with real-world datasets for external validity. Experiments show that performance rankings vary significantly across SCM families and protocols, indicating that strong results on one benchmark do not generalize to others, especially in the context of causal discovery foundation models.