arXiv:2607. 08349v1 Announce Type: new Abstract: Mechanistic interpretability often evaluates explanations by intervening on a model: swapping hidden states, patching activations, ablating components, or comparing a compressed model to the original one.
By Amir Asiaee
arXiv:2607. 28308v1 Announce Type: new Abstract: Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions.
By Huiyuan Tian, Bonan Xu, Shijian Li
Sparse mixture-of-experts (MoE) language models route each token to multiple experts, suggesting a geometric account of their benefit: co-selected experts should contribute distinct representation directions. Existing evidence often conflates route coherence, candidate quality, and candidate-by-context interaction.
The paper investigates whether the factors highlighted by large language models (LLMs) as most influential on their decisions truly reflect necessity or sufficiency in influencing outcomes. By applying controlled black‑box interventions across eight models from Claude, GPT, and Gemini, the authors quantify necessity and sufficiency scores for each factor and compare them to the models’ self‑reported top three factors. Results show modest correlations (≈0.35–0.58) and reveal that the cited top factors often fail to capture the strongest measured influences, indicating limitations in current explanation practices.
By Urja Pawar, Rajitha Ramanayake, Nabeel Kemal, Ashwin Kandath, Owen O'Neill, Guillaume Bourgeon, Houssem Chatbri
arXiv:2606. 25092v1 Announce Type: new Abstract: Sparse Mixture-of-Experts (MoE) models route each token to a few of many experts, inviting the hypothesis that experts form functional modules tied to capabilities or languages.
By Tony Salomone, Deep Gandhi, Ali Asaria
arXiv:2606. 27510v1 Announce Type: new Abstract: Activation patching is the primary tool in mechanistic interpretability.
By Sankaran Vaidyanathan, David Arbour, Aaron Mueller, Scott Niekum, David Jensen