arXiv:2609.18961v1 Announce Type: new
Abstract: Mechanistic interpretability identifies sparse subsets of heads and MLP blocks that carry specific behaviors. We ask whether such causal signals can gu...
By Son Ha Xuan, Phat T. Tran-Truong, Xuan-Bach Le
arXiv:2606. 23880v1 Announce Type: new Abstract: From climate teleconnections to gene regulation, modern time-series datasets encompass tens or hundreds of interacting variables, making causal discovery increasingly challenging.
By Mohammad Fesanghary, Abhinav Havaldar
arXiv:2607. 05806v1 Announce Type: new Abstract: Training data for machine learning is routinely collected by a selection process the model never sees: loans are observed only when granted, outcomes only when a test was ordered.
By Gunner Levi Howe
arXiv:2608. 15725v1 Announce Type: new Abstract: Predictive models in clinical and regulated settings must be accurate and fully auditable.
By Srikumar Krishnamoorthy
The paper introduces SAEScientist-Bench, a benchmark that tests whether AI agents can autonomously conduct mechanistic interpretability research using Sparse Autoencoders (SAEs). Agents are tasked with designing contrastive probes and navigating a large feature dictionary in Gemma-2-9B-IT to identify optimal features for a target concept, with performance measured against expert-curated references on activation rank, concept selectivity, and causal steering. Results show that while frontier agents can discover features and outperform controls, they still lag behind expert baselines, especially in causal steering, highlighting both the potential and current limitations of closed-loop autonomous AI research.
By Yuqiao Tan, Shizhu He, Jun Zhao, Kang Liu
The paper introduces CELLAUDIT, a method for auditing whether inputs claimed to influence predictive models actually do so. By testing if an input can enter the computation, whether predictions depend on it, and if that dependence improves observed responses, the authors evaluate agent-generated predictors on a morphology‑transcriptomics benchmark (BBBC047). Their findings show that many models claim compound contributions that are not supported by the data, and that falsification‑guided revisions can recover genuine input effects while improving performance.
By Mengran Li, Bo Li, Chengyang Zhang, Yang Yan, Jinfeng Xu, Zhenchao Tang
arXiv:2608. 01023v1 Announce Type: new Abstract: We present Caliber, an output-perturbation defense against model extraction that formulates noise selection as a calibration problem: how much the defense degrades the supervision signal used to train a surrogate, and the provable per-input query cost of recovering the clean logits.
By Chi Wang, Hanwen Wang, Yu Xia, Zihan Wang, Guangdong Bai
arXiv:2608. 07914v1 Announce Type: new Abstract: Behavioral contamination detectors can return "no evidence" either because a benchmark is clean or because the audit has little power.
By Ibne Farabi Shihab, Sanjeda Akter, Anuj Sharma
The paper investigates mechanistic interpretability, focusing on how automated circuit discovery is evaluated. It shows that the commonly used faithfulness objective can favor circuits that reproduce a model’s behavior poorly, creating an objective-level recovery gap. Experiments on four human-reference tasks and InterpBench reveal that many discovery methods misrank candidate circuits, and that restoring excluded signals can correct most of these misrankings without altering the circuits’ behavior.
By Chuqin Geng, Li Zhang, Haolin Ye, Mark Zhang, Luke Zhang, Xujie Si
arXiv:2607. 25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time.
By Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal, Dipesh Mahato
The paper proposes a flexible four‑coefficient parameterization for per‑token gating in on‑policy knowledge distillation, unifying existing methods such as EOPD and ToDi as special cases. Experiments on TweetEval with a Qwen3 teacher‑student pair show that the full family of gating configurations outperforms single‑channel baselines in most cells, and dynamic gating beats static baselines in a majority of isolated comparisons. The authors present the framework mainly as a shared coordinate system for comparing gating designs rather than as definitive statistical evidence.
By Suwan Wu, Yumeng Lin, Pengcheng Yuan, Xiaolong Jiang
arXiv:2607. 19364v2 Announce Type: replace Abstract: Activation steering adds a residual-stream direction at inference time, providing lightweight behavioral control without fine-tuning.
By Oshayer Siddique, J. M Areeb Uzair Alam, Md Jobayer Rahman Rafy, Syed Rifat Raiyan, Hasan Mahmud, Md Kamrul Hasan