arXiv Machine Learning By Xianzhi Zeng, Jiangneng Li, Gao Cong

To Explore The Strange New World Beyond Data Distribution: System Behavior, Causality Tax, and Non-causal Base Model

Read the original on arXiv Machine Learning →

The paper argues that the causality of language models may be unnecessary or suboptimal when system behavior—extra dominant factors beyond data distribution—is treated as a first‑principle Bayesian feature. It introduces the SBD framework, incorporating system behavior into the evidence lower bound, and demonstrates a counter‑intuitive Causality Tax where ignoring these factors leads to structural error. Using a non‑causal variational family called Green Shell, the authors show through theoretical bounds, implicit measurements, and Neural Tangent Kernel analysis that this approach yields tighter error bounds and improved generalization compared to causal models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

ProximalFM: Amortized Proximal Causal Inference under Hidden Confounding

ProximalFM is a transformer‑based model that uses prior‑data fitted networks (PFNs) to perform Bayesian proximal causal inference under hidden confounding. By training on synthetic data generated from structural causal models with oracle counterfactuals, it amortizes the Bayesian operator inversion into a single forward pass, producing posterior estimates of the conditional average treatment effect (CATE). The approach consistently outperforms prior methods across various proximal regimes, especially when latent confounding is strong and proxy variables are weakly informative, and it requires no dataset‑specific tuning.

By Christophe Muller, Ayub Kharel, Alex Luedtke, Chan Park, Eric Tchetgen Tchetgen, Juan L. Gamella, Rahul Krishnan, Ricardo Silva, Jakob Zeitler