arXiv:2606. 27510v1 Announce Type: new Abstract: Activation patching is the primary tool in mechanistic interpretability.
By Sankaran Vaidyanathan, David Arbour, Aaron Mueller, Scott Niekum, David Jensen
arXiv:2607. 08349v1 Announce Type: new Abstract: Mechanistic interpretability often evaluates explanations by intervening on a model: swapping hidden states, patching activations, ablating components, or comparing a compressed model to the original one.
By Amir Asiaee
arXiv:2401. 04890v2 Announce Type: replace-cross Abstract: This work introduces a novel principle for disentanglement we call mechanism sparsity regularization, which applies when the latent factors of interest depend sparsely on observed auxiliary variables and/or past latent factors.
By S\'ebastien Lachapelle, Pau Rodr\'iguez L\'opez, Yash Sharma, Katie Everett, R\'emi Le Priol, Alexandre Lacoste, Simon Lacoste-Julien
arXiv:2507. 14661v2 Announce Type: replace-cross Abstract: Semi-supervised domain adaptation (SSDA) seeks to achieve accurate predictions in a target domain with limited labeled target data by exploiting abundant source and unlabeled target data.
By Wooseok Ha, Yuansi Chen
arXiv:2606. 19610v1 Announce Type: cross Abstract: Recent work on Kan-Do-Calculus (KDC) has established that the boundary between passive observation and active intervention in causal inference is a category-theoretic bi-adjunction, with interventions modeled by left Kan extensions and conditioning by right Kan extensions.
By Sridhar Mahadevan
arXiv:2607. 03999v1 Announce Type: cross Abstract: Estimating heterogeneous treatment effects (CATE) requires simultaneously detecting effect modification and quantifying estimation uncertainty.
By Pantelis Z. Hadjipantelis, Josephine Chiang, Karthik Nagesh