arXiv:2606. 17516v1 Announce Type: cross Abstract: Causal discovery from observational data remains challenging due to the need to recover directed structure and latent confounding without interventions.
By Patrick Bl\"obaum, Krishnakumar Balasubramanian, Shiva Prasad Kasiviswanathan
Matryoshka Attribution (MAttr) is a mask‑learning method that identifies nested subsets of a language model’s internal components by minimizing downstream loss. It uses a differentiable sigmoid top‑k operator and randomizes sparsity during training to produce an attribution ordering of components. MAttr tops the Mechanistic Interpretability Benchmark leaderboard and can be applied via reinforcement learning to pinpoint weight changes that control behaviors such as refusal in Llama 3.1 8B Instruct, where restoring just 1% of weights removes refusals while preserving capabilities.
By Aryaman Arora, Kirill Acharya, Nathan Hu, Yanzhe Zhang, Noah Goodman, Dan Jurafsky, Christopher Potts
arXiv:2603. 24304v2 Announce Type: replace-cross Abstract: Graph Neural Networks (GNNs) deliver strong performance on graph tasks, but their accuracy drops significantly under out-of-distribution (OOD) scenarios.
By Bowen Lu, Liangqiang Yang, Teng Li, Kun Zhang
arXiv:2609.35890v1 Announce Type: new
Abstract: Sparse-autoencoder decomposability and concentrated feature attribution are increasingly treated as evidence that a model's computation is easier to re...
By Adam Elimadi
arXiv:2601. 21996v2 Announce Type: replace-cross Abstract: While Mechanistic Interpretability has identified interpretable circuits in LLMs, their causal origins in training data remain elusive.
By Jianhui Chen, Yuzhang Luo, Liangming Pan
arXiv:2609.22566v1 Announce Type: cross
Abstract: Knowledge distillation (KD) aims to compress high-performance teacher LLMs into lightweight students. However, distilled students often exhibit subst...
By Dileesha Kannangara, Sanghamitra Dutta
arXiv:2507. 18043v2 Announce Type: replace-cross Abstract: Inference-time steering methods offer a lightweight alternative to fine-tuning large language models (LLMs) and vision-language models (VLMs) by modifying internal activations at test time without updating model weights.
By Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal
arXiv:2602. 07008v3 Announce Type: replace-cross Abstract: Reliable models should not only predict correctly, but also justify decisions with acceptable evidence.
By Ruoyu Chen, Shangquan Sun, Xiaoqing Guo, Sanyi Zhang, Kangwei Liu, Shiming Liu, Zhangcheng Wang, Qunli Zhang, Wei Wang, Hua Zhang, Xiaochun Cao
arXiv:2606. 23872v1 Announce Type: cross Abstract: As generative models increasingly produce samples that are indistinguishable from human-created content, it becomes difficult to determine whether a given data point was part of a model's natural training set or was generated by the model itself, especially when models memorize and reproduce training data.
By Bihe Zhao, Michel Meintz, Juangui Xu, Franziska Boenisch, Adam Dziedzic
arXiv:2603. 01372v2 Announce Type: replace-cross Abstract: Concept Bottleneck Models (CBMs) enhance the interpretability of end-to-end neural networks by introducing a layer of concepts and predicting the class label from the concept predictions.
By Weixin Chen, Han Zhao
arXiv:2606. 27114v1 Announce Type: new Abstract: Uplift modeling, crucial for estimating individual treatment effects (ITE), faces dual challenges: flexibly leveraging inter-group similarity to enhance discriminative power and debiasing under unobserved confounding scenarios.
By Haoran Zhang, Chuanpu Li, Yuxin Fu, Bin Tong, Guan Wang, Bo Zheng, Feng Zhou
arXiv:2602.07008v5 Announce Type: replace
Abstract: Reliable models should not only predict correctly, but also base their decisions on acceptable evidence. However, conventional supervised learning...
By Ruoyu Chen, Shangquan Sun, Xiaoqing Guo, Kangwei Liu, Sanyi Zhang, Zhangcheng Wang, Shiming Liu, Qunli Zhang, Wei Wang, Hua Zhang, Xiaochun Cao