arXiv Machine Learning

When few labeled target data suffice: a theory of semi-supervised domain adaptation via fine-tuning from multiple adaptive starts

arXiv:2507. 14661v2 Announce Type: replace-cross Abstract: Semi-supervised domain adaptation (SSDA) seeks to achieve accurate predictions in a target domain with limited labeled target data by exploiting abundant source and unlabeled target data.

arXiv Machine Learning
2d ago

CIDER-FM: Foundation Models for Causal Inference from Diverse Experimental Regimes

CIDER-FM is a causal foundation model that combines finite observational data with surrogate-interventional datasets to predict target conditional interventional distributions more accurately than using observational data alone. It employs an intervention-aware representation and hierarchical three‑axis attention to integrate information across variables, samples, and experimental regimes. Experiments on synthetic graphs, simulated data, and real‑world Causal Chambers data show that incorporating experimental context improves CID prediction performance.

By Yuche Gao, Arik Reuter, Siyuan Guo, Anish Dhir, Bernhard Sch\"olkopf, Adrian Weller
arXiv Machine Learning
Aug 19

TabCausal: Pretraining Across Causal Environments for Tabular Causal Discovery

TabCausal is a causal discovery foundation model that learns to map datasets directly to causal graphs by pretraining across diverse causal environments. It uses a dynamic task construction strategy to expose the model to varied graph priors, mechanisms, noise models, dimensions, sample sizes, and intervention regimes, improving transferability from observational and mixed‑interventional data. On large synthetic benchmarks and a new protocol‑guided semantic benchmark, TabCausal outperforms many classical baselines and shows robust structure recovery, especially when interventional evidence is available.

By Zi-Rong Li, Si-Yang Liu, Tian-Zuo Wang, Han-Jia Ye
arXiv Machine Learning
Jun 18

Anti-causal domain generalization: Leveraging unlabeled data

arXiv:2602. 17187v2 Announce Type: replace-cross Abstract: The problem of domain generalization concerns learning predictive models that are robust to distribution shifts when deployed in new, previously unseen environments.

By Sorawit Saengkyongam, Juan L. Gamella, Andrew C. Miller, Jonas Peters, Nicolai Meinshausen, Christina Heinze-Deml
arXiv AI
Jul 14

CDFM: Towards a General-Purpose Causal Discovery Foundation Model

arXiv:2607. 11508v1 Announce Type: cross Abstract: Causal discovery, the process of recovering underlying causal structures from observational data, is a fundamental pursuit across scientific disciplines.

By Jie Qiao, Ruichu Cai, Zijian Li, Weilin Chen, Pengfei Hua, Boyan Xu, Zhengming Chen, Zhifeng Hao, Peng Cui