arXiv:2607. 08349v1 Announce Type: new Abstract: Mechanistic interpretability often evaluates explanations by intervening on a model: swapping hidden states, patching activations, ablating components, or comparing a compressed model to the original one.
By Amir Asiaee
arXiv:2607. 28308v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions.
By Huiyuan Tian, Bonan Xu, Shijian Li
Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often conflates route coherence, candidate quality, and candidate-by-context interaction.
The paper investigates whether the factors highlighted by large language models (LLMs) as most influential on their decisions truly reflect necessity or sufficiency in influencing outcomes. By applying controlled black‑box interventions across eight models from Claude, GPT, and Gemini, the authors quantify necessity and sufficiency scores for each factor and compare them to the models’ self‑reported top three factors. Results show modest correlations (≈0.35–0.58) and reveal that the cited top factors often fail to capture the strongest measured influences, indicating limitations in current explanation practices.
By Urja Pawar, Rajitha Ramanayake, Nabeel Kemal, Ashwin Kandath, Owen O'Neill, Guillaume Bourgeon, Houssem Chatbri
arXiv:2606. 25092v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) models route each token to a few of many experts, inviting the hypothesis that experts form functional modules tied to capabilities or languages.
By Tony Salomone, Deep Gandhi, Ali Asaria
arXiv:2606. 27510v1 Announce Type: new Abstract: Activation patching is the primary tool in mechanistic interpretability.
By Sankaran Vaidyanathan, David Arbour, Aaron Mueller, Scott Niekum, David Jensen
arXiv:2606. 03780v1 Announce Type: cross Abstract: Causal tracing of factual recall has been studied predominantly in dense transformer language models, where interventions localize information flow to layers or feed-forward modules.
By Yuetian Lu, Ali Modarressi, Yihong Liu, Hinrich Sch\"utze
arXiv:2608. 12555v1 Announce Type: new Abstract: Predictive explanation methods attribute a model output; they do not, by themselves, attribute an intervention effect on the real-world outcome.
By Michael Georgiades, Charalambia Varnava
The paper introduces GeoACE, a five‑expert framework for estimating heterogeneous treatment effects that blends a common anchor‑correction estimator with overlap‑aware and outcome‑guided geometries. The ensemble’s task‑level weights are learned from internal validation predictions, frozen before test evaluation, and applied to experts refitted on the full development data. Adding the outcome‑free, overlap‑aware expert O‑Phi‑ACE consistently improves performance across seven benchmarks, achieving the lowest average rank among 11 comparators.
By Ali Haghpanah Jahromi, Mohammad Taheri
arXiv:2604. 07650v2 Announce Type: replace Abstract: The rapid growth of the large language model (LLM) ecosystem raises a critical question: are seemingly diverse models truly independent?
By Chenchen Kuai, Jiwan Jiang, Zihao Zhu, Hao Wang, Keshu Wu, Zihao Li, Yunlong Zhang, Chenxi Liu, Zhengzhong Tu, Zhiwen Fan, Yang Zhou
The paper investigates how the definition of influence—specifically the behavior being attributed, the intervention on training data, and the counterfactual training process—affects rankings produced by influence estimators. It formalizes influence as a counterfactual estimand, distinguishes specification mismatch from approximation error, and categorizes existing estimators by their implied specifications. Experiments demonstrate that different specifications can lead to markedly different rankings, and that careful specification choice improves attribution quality in tasks such as noisy label detection and large‑language‑model attribution.
By Zhe Li, Wei Zhao, Peixin Zhang, Jun Sun
arXiv:2608. 11212v1 Announce Type: new Abstract: Top-k Mixture-of-Experts (MoE) routing is discontinuous, so a deployment-motivated numerical disturbance -- simulated 4-bit KV-cache quantization read by a protected BF16 gate -- pushes tokens across decision boundaries and flips which experts fire.
By Parvel Gu