arXiv AI By Jinyuan Zhang, Peng He, Yin Yuan, He Hu, ShengShuo Jiao

HiPACE: Hierarchical Phase-Boundary Analysis and Controlled Evaluation of Feature Absorption in Sparse Autoencoders

Read the original on arXiv AI →

The paper introduces HiPACE, a protocol for evaluating feature absorption in sparse autoencoders (SAEs). It derives a closed‑form phase boundary λ_c(k,α)=α^2k/(k-1) that predicts when a parent concept will dominate over its children in the SAE dictionary. Experiments on synthetic data and Pythia‑160m SAEs confirm the boundary’s sharpness and demonstrate causal effects of family‑direction activations on parent‑category logits.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 15

From Geometric Recovery to Causal Validation: A Reproducible Audit of Sparse Autoencoder Features, from Superposition Geometry to Causal Inertness

arXiv:2607. 12166v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are the standard for decomposing superposed neural representations into interpretable features, and evaluation relies predominantly on correlational recovery metrics -- cosine similarity between ground-truth directions and decoder atoms.

By Mohamed Abdessalem Bal
arXiv Machine Learning
Aug 4

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations

arXiv:2605. 28149v2 Announce Type: replace Abstract: Sparse Autoencoders (SAEs) extract interpretable features from Large Language Model activations, but standard variants enforce non-negative latents, so a bidirectional semantic axis (e.

By Bartosz Wieciech, Zmnako Awrahman, Marcin Czelej, Victor Hugo Jaramillo Velasquez, Wioletta Stobieniecka
arXiv AI
Sep 17

Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition

Linear probes can decode safety‑relevant concepts such as truthfulness from language‑model activations, but probe accuracy may reflect only decodability, not causal influence on model behavior. The authors show that probe weight geometry alone cannot identify the features the model actually uses, because geometrically aligned features need not be causally relevant. They introduce a sparse‑autoencoder (SAE) decomposition that ranks features by probe alignment and gradient sensitivity, and demonstrate that ablating shared, probe‑only, and random feature sets reveals a sharp dissociation: shared features drive model output changes far more than probe‑only or random features, confirming that causal relevance requires intervention beyond weight geometry.

By Devesh Tiwari, Camille Davis, Shivank Sinha, Talia Weaver, Aditya Shah, Maheep Chaudhary