Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition
Read the original on arXiv AI →Linear probes can decode safety‑relevant concepts such as truthfulness from language‑model activations, but probe accuracy may reflect only decodability, not causal influence on model behavior. The authors show that probe weight geometry alone cannot identify the features the model actually uses, because geometrically aligned features need not be causally relevant. They introduce a sparse‑autoencoder (SAE) decomposition that ranks features by probe alignment and gradient sensitivity, and demonstrate that ablating shared, probe‑only, and random feature sets reveals a sharp dissociation: shared features drive model output changes far more than probe‑only or random features, confirming that causal relevance requires intervention beyond weight geometry.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.