TabCausal is a causal discovery foundation model that learns to map datasets directly to causal graphs by pretraining across diverse causal environments. It uses a dynamic task construction strategy to expose the model to varied graph priors, mechanisms, noise models, dimensions, sample sizes, and intervention regimes, improving transferability from observational and mixed‑interventional data. On large synthetic benchmarks and a new protocol‑guided semantic benchmark, TabCausal outperforms many classical baselines and shows robust structure recovery, especially when interventional evidence is available.
By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
arXiv:2603. 10254v2 Announce Type: replace Abstract: Synthetic tabular data generation addresses data scarcity and privacy constraints in a variety of domains.
By Davide Tugnoli, Andrea De Lorenzo, Marco Virgolin, Giovanni Cin\`a
The paper investigates how synthetic pretraining priors used in tabular foundation models (TFMs) influence downstream performance. By reconstructing the synthetic data generators of four TFMs and comparing their generated tasks to two popular tabular benchmarks using structural descriptors, the authors measure structural coverage and normalized density. They find that some generators provide broader and denser support for benchmark tasks, and that stronger synthetic-to-benchmark support generally correlates with better model performance.
By He Zhao, Ryan Thompson, Daniel M. Steinberg, Ashfaqur Rahman, Edwin V. Bonilla, Cheng Soon Ong
CausalArena is a new benchmark designed to evaluate causal discovery methods in the era of foundation models. It unifies synthetic structural causal models (SCMs), semantically grounded SCMs, and formula‑grounded SCMs, while also including real‑world datasets for external validation. Experiments show that performance rankings vary widely across different SCM families and protocols, indicating that strong results on one benchmark do not necessarily transfer to others.
By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
arXiv:2603. 10823v2 Announce Type: replace-cross Abstract: Deep generative models can help with data scarcity and privacy by producing synthetic training data, but they struggle in low-data, imbalanced tabular settings to fully learn the complex data distribution.
By Xiaofeng Lin, Seungbae Kim, Zhuoya Li, Zachary DeSoto, Charles Fleming, Guang Cheng
CausalArena is a unified, evolvable benchmark designed to evaluate causal discovery methods across diverse structural causal models (SCMs). It incorporates synthetic SCMs for controlled structural variation, semantic operational SCMs for human-auditable environments, and formula-grounded SCMs to test discovery under explicit scientific mechanisms, along with real-world datasets for external validity. Experiments show that performance rankings vary significantly across SCM families and protocols, indicating that strong results on one benchmark do not generalize to others, especially in the context of causal discovery foundation models.
GENSCRIPT is an inference‑only pipeline that generates synthetic data without training a generative model. It creates a deterministic statistical profile of the source data, uses a language model to infer field semantics and cross‑column constraints, and then compiles these into an executable sampler that works for single‑table, temporal, and relational data. The method builds generators in minutes, samples large datasets quickly, and achieves fidelity comparable to leading methods while preserving key data relationships such as 1‑to‑1 mappings and primary‑foreign key constraints.
By Zilong Zhao, Abdul Raheem, Jiayu Li, Sohei Arisaka, Darius Lim Hong Yi, Milad Abdollahzadeh, Uzair Javaid, Biplab Sikdar
arXiv:2510.24046v2 Announce Type: replace-cross
Abstract: Existing tabular data generation methods primarily focus on matching statistical distributions between real and synthetic data, often overloo...
By Tu Anh Hoang Nguyen, Dang Nguyen, Tri-Nhan Vo, Thuc Duy Le, Trung Le, Sunil Gupta
CausalProfiler is a synthetic benchmark generator designed to evaluate causal machine learning (Causal ML) methods more rigorously and transparently. It randomly samples causal models, data, queries, and ground truths based on explicit design choices across observation, intervention, and counterfactual reasoning levels, providing coverage guarantees and transparent assumptions. The authors demonstrate its utility by testing several state‑of‑the‑art methods under diverse conditions, both within and outside the identification regime, highlighting the insights CausalProfiler can reveal.
By Panayiotis Panayiotou, Audrey Poinsot, Alessandro Leite, Nicolas Chesneau, Marc Schoenauer, \"Ozg\"ur \c{S}im\c{s}ek
arXiv:2607. 11508v1 Announce Type: cross Abstract: Causal discovery, the process of recovering underlying causal structures from observational data, is a fundamental pursuit across scientific disciplines.
By Jie Qiao, Ruichu Cai, Zijian Li, Weilin Chen, Pengfei Hua, Boyan Xu, Zhengming Chen, Zhifeng Hao, Peng Cui
arXiv:2609.36881v1 Announce Type: new
Abstract: Causal foundation models (CFMs) pre-trained on data generated from various structural causal models (SCMs) have been proposed for estimating causal eff...
By Heejin Jung, Gyeongdeok Seo, Hoyoon Byun, Joseph Lee, Kyungwoo Song
arXiv:2607. 08093v1 Announce Type: new Abstract: Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use.
By Andrej Leban, Yuekai Sun