The paper introduces a causal mediation framework to separate direct discrimination from structural inequality in AI-driven credit decisions. Using Pearl’s natural direct and indirect effects, it presents an identification strategy under treatment‑induced confounding and proposes a doubly‑robust estimator with efficiency guarantees. Empirical analysis of 89,465 mortgage applications shows that about 77% of racial denial disparities stem from financial mediators, while the remaining 23% represents a conservative lower bound on direct discrimination.
By Duraimurugan Rajamanickam
arXiv:2606. 10703v1 Announce Type: new Abstract: Interpretability methods routinely use population-level summary statistics over observed model behaviour to license claims about the effects of targeted interventions on specific computations; in Pearl's terms, they treat rung-1 associational evidence as if it supported rung-2 interventional conclusions, a move whose validity is rarely tested.
By Leonard Engmann, Christian Medeiros Adriano, Holger Giese
arXiv:2609.36881v1 Announce Type: new
Abstract: Causal foundation models (CFMs) pre-trained on data generated from various structural causal models (SCMs) have been proposed for estimating causal eff...
By Heejin Jung, Gyeongdeok Seo, Hoyoon Byun, Joseph Lee, Kyungwoo Song
arXiv:2607. 29484v1 Announce Type: cross Abstract: Interventional data is widely regarded as the gold standard for teaching models causal reasoning.
By Xining Xun
arXiv:2606. 08275v1 Announce Type: cross Abstract: When an LLM agent fails -- issues a refund it should not have, calls the wrong tool, leaks data -- existing tooling answers what happened (observability) or whether it passed (evaluation), but not which step caused the failure.
By Jaineet Shah
The paper proposes Causal Evidentiary Governance (CEG), a framework that requires regulated institutions to maintain a versioned directed acyclic graph (DAG) separating allowable from disallowed causal pathways in high‑risk machine learning systems. CEG introduces the Causal Harm Rate to quantify prediction variation due to disallowed pathways and pairs each decision with a signed Decision‑Evidence Packet (DEP) that cryptographically links the prediction to the DAG and path‑specific attributions, enabling efficient inclusion proofs via a Merkle tree. Empirical validation on synthetic credit data and the German Credit dataset demonstrates that CEG more clearly isolates causal effects than traditional fairness metrics and that a proof‑of‑concept implementation shows operational feasibility with manageable performance tradeoffs.
By Samah Kareem, Bar{\i}\c{s} \c{C}elikta\c{s}
The paper introduces GeoACE, a five‑expert framework for estimating heterogeneous treatment effects that blends a common anchor‑correction estimator with overlap‑aware and outcome‑guided geometries. The ensemble’s task‑level weights are learned from internal validation predictions, frozen before test evaluation, and applied to experts refitted on the full development data. Adding the outcome‑free, overlap‑aware expert O‑Phi‑ACE consistently improves performance across seven benchmarks, achieving the lowest average rank among 11 comparators.
By Ali Haghpanah Jahromi, Mohammad Taheri
arXiv:2606. 19625v2 Announce Type: replace-cross Abstract: We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social-reasoning versus STEM-reasoning in OLMo3-7B.
By Glenn Matlin, Chandreyi Chakraborty, Saehee Eom, Mika Okamoto, Rayan Castilla, Louis Jaburi, Alvin Deng, Taywon Min, Lucia Quirke, Stella Biderman, Mark Riedl
The paper introduces Actionable Case-Based Feature Importance (A‑CBFI), a framework that integrates structural causal models with counterfactual recourse for tabular machine learning. A‑CBFI isolates synergistic interaction bottlenecks and releases suppressive structural locks, concentrating over 98.3% of intervention effort on diagnosed root causes. Empirical tests in finance and healthcare show a 76.9% reduction in active human intervention while keeping recourse costs comparable to exhaustive causal methods.
By Sejong Oh
arXiv:2609.07944v1 Announce Type: new
Abstract: Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recove...
By Yonghong Zhang, Ricardo Correia, Isabel M. Parra, Yong Xie
arXiv:2608. 20187v1 Announce Type: cross Abstract: Practitioners inferring causality from observational data usually rely on a single method and treat its output as causal truth.
By Manish Gupta, Dipanjan De
arXiv:2606. 27114v1 Announce Type: new Abstract: Uplift modeling, crucial for estimating individual treatment effects (ITE), faces dual challenges: flexibly leveraging inter-group similarity to enhance discriminative power and debiasing under unobserved confounding scenarios.
By Haoran Zhang, Chuanpu Li, Yuxin Fu, Bin Tong, Guan Wang, Bo Zheng, Feng Zhou