arXiv:2609.07061v2 Announce Type: replace
Abstract: Physics-Informed Neural Networks (PINNs) embed PDE residuals into neural network training, but their internal representations remain opaque: it is...
By Nandita N. Patil, Eshwar R. A., Gajanan V. Honnavar
arXiv:2609.07061v1 Announce Type: new
Abstract: Physics-Informed Neural Networks (PINNs) embed PDE residuals into neural network training, but their internal representations remain opaque: it is unkn...
By Nandita N. Patil, Eshwar R. A., Gajanan V. Honnavar
Linear probes can decode safety‑relevant concepts such as truthfulness from language‑model activations, but probe accuracy may reflect only decodability, not causal influence on model behavior. The authors show that probe weight geometry alone cannot identify the features the model actually uses, because geometrically aligned features need not be causally relevant. They introduce a sparse‑autoencoder (SAE) decomposition that ranks features by probe alignment and gradient sensitivity, and demonstrate that ablating shared, probe‑only, and random feature sets reveals a sharp dissociation: shared features drive model output changes far more than probe‑only or random features, confirming that causal relevance requires intervention beyond weight geometry.
By Devesh Tiwari, Camille Davis, Shivank Sinha, Talia Weaver, Aditya Shah, Maheep Chaudhary
arXiv:2609. 12591v1 Announce Type: new Abstract: Foundation models are increasingly adapted through fine-tuning, model editing, and alignment procedures while retaining previously acquired capabilities.
By Hendrik Droste, Christian Medeiros Adriano, Kathrin Korte, Holger Giese
arXiv:2606. 18383v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are increasingly used to extract interpretable features from language models (LMs), yet a central question remains: when can an SAE-based explanation be treated as a faithful view of an underlying frozen LM We study this through a post-hoc generalization framework that certifies the LM via a sparse proxy, obtained by replacing a native hidden activation with its pretrained SAE reconstruction.
By Dibyanayan Bandyopadhyay, Asif Ekbal
arXiv:2606. 12138v1 Announce Type: cross Abstract: Sparse autoencoders (SAEs) are widely used to interpret neural network representations, but their utility depends on whether the learned features are reproducible across training runs.
By Gleb Gerasimov, Timofei Rusalev, Nikita Balagansky, Daniil Laptev, Vadim Kurochkin, Daniil Gavrilov
The paper introduces HiPACE, a protocol for evaluating feature absorption in sparse autoencoders (SAEs). It derives a closed‑form phase boundary λ_c(k,α)=α^2k/(k-1) that predicts when a parent concept will dominate over its children in the SAE dictionary. Experiments on synthetic data and Pythia‑160m SAEs confirm the boundary’s sharpness and demonstrate causal effects of family‑direction activations on parent‑category logits.
By Jinyuan Zhang, Peng He, Yin Yuan, He Hu, ShengShuo Jiao
arXiv:2605. 28149v2 Announce Type: replace Abstract: Sparse Autoencoders (SAEs) extract interpretable features from Large Language Model activations, but standard variants enforce non-negative latents, so a bidirectional semantic axis (e.
By Bartosz Wieciech, Zmnako Awrahman, Marcin Czelej, Victor Hugo Jaramillo Velasquez, Wioletta Stobieniecka
arXiv:2607. 20596v1 Announce Type: new Abstract: Sparse autoencoder (SAE) features are used to interpret and steer large language models, yet whether a feature's causal role is stable across SAE families remains untested.
By Seonglae Cho, Zekun Wu, Kleyton Da Costa, Rishi Kalra, Ilham Wicaksono, Adriano Koshiyama
arXiv:2511. 09432v2 Announce Type: replace Abstract: Machine learning (ML) models achieve remarkable performance but remain hard to interpret due to their scale and complexity.
By Ege Erdogan, Ana Lucic
arXiv:2607. 19618v1 Announce Type: cross Abstract: Genomic language models achieve strong performance across regulatory-genomics tasks, yet what these models internally represent remains opaque, and the field lacks a principled procedure for verifying that an apparent ``concept'' inside a model is real rather than an artifact of sequence composition.
By Sarwan Ali
arXiv:2508. 16560v4 Announce Type: replace-cross Abstract: Sparse Autoencoders (SAEs) extract features from LLM internal activations, meant to correspond to interpretable concepts.
By David Chanin, Adri\`a Garriga-Alonso