The Curse of Multiple Mediators: Hidden Interaction Effects in Activation Patching
arXiv:2606. 27510v1 Announce Type: new Abstract: Activation patching is the primary tool in mechanistic interpretability.
arXiv:2606. 19410v1 Announce Type: cross Abstract: Signed pairwise interaction scores fundamentally conflate uniqueness (U), redundancy (R), and synergy (S).
arXiv:2606. 27510v1 Announce Type: new Abstract: Activation patching is the primary tool in mechanistic interpretability.
arXiv:2607. 08349v1 Announce Type: new Abstract: Mechanistic interpretability often evaluates explanations by intervening on a model: swapping hidden states, patching activations, ablating components, or comparing a compressed model to the original one.
arXiv:2401. 04890v2 Announce Type: replace-cross Abstract: This work introduces a novel principle for disentanglement we call mechanism sparsity regularization, which applies when the latent factors of interest depend sparsely on observed auxiliary variables and/or past latent factors.
arXiv:2507. 14661v2 Announce Type: replace-cross Abstract: Semi-supervised domain adaptation (SSDA) seeks to achieve accurate predictions in a target domain with limited labeled target data by exploiting abundant source and unlabeled target data.
arXiv:2606. 19610v1 Announce Type: cross Abstract: Recent work on Kan-Do-Calculus (KDC) has established that the boundary between passive observation and active intervention in causal inference is a category-theoretic bi-adjunction, with interventions modeled by left Kan extensions and conditioning by right Kan extensions.
arXiv:2607. 03999v1 Announce Type: cross Abstract: Estimating heterogeneous treatment effects (CATE) requires simultaneously detecting effect modification and quantifying estimation uncertainty.
arXiv:2603. 13326v2 Announce Type: replace-cross Abstract: Multimodal Transformers often produce predictions without clarifying how different modalities jointly support a decision.
arXiv:2606. 17516v1 Announce Type: cross Abstract: Causal discovery from observational data remains challenging due to the need to recover directed structure and latent confounding without interventions.
arXiv:2608.28991v1 Announce Type: cross Abstract: Causal representation learning (CRL) aims to recover latent causal variables and their structural relations from high-dimensional observations. Exist...
arXiv:2609.35890v1 Announce Type: new Abstract: Sparse-autoencoder decomposability and concentrated feature attribution are increasingly treated as evidence that a model's computation is easier to re...
The paper introduces Actionable Case-Based Feature Importance (A‑CBFI), a framework that integrates structural causal models with counterfactual recourse for tabular machine learning. A‑CBFI isolates synergistic interaction bottlenecks and releases suppressive structural locks, concentrating over 98.3% of intervention effort on diagnosed root causes. Empirical tests in finance and healthcare show a 76.9% reduction in active human intervention while keeping recourse costs comparable to exhaustive causal methods.
arXiv:2603. 02204v2 Announce Type: replace Abstract: Selective conformal prediction can yield substantially tighter uncertainty sets when we can identify calibration examples that are exchangeable with the test example.