Universal Activation Verbalizer: A Unified Framework for Cross-Model Activation Explanation
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
Sparse autoencoders (SAEs) are proposed to extract numerous features from large language model (LLM) representations, yet explaining these features still relies primarily on external observation. This reliance leads to superficial explanations inferred from observed model behavior and computational inefficiency from collecting such behavioral evidence at scale.
arXiv:2608. 09521v1 Announce Type: new Abstract: Activation-based tools are usually tied to one model's native hidden space, requiring probes, sparse autoencoders, and natural-language interpreters to be rebuilt or rediscovered for each new language model.
arXiv:2606. 16137v1 Announce Type: cross Abstract: Speech deepfake detection (SDD) systems require trustworthy explanations for reliable decision-making.
arXiv:2508.08966v2 Announce Type: replace Abstract: The attention mechanism lies at the core of the transformer architecture, providing an interpretable model-internal signal that has motivated a gro...
SharedSAE demonstrates that a single sparse autoencoder can replace multiple model‑specific SAEs by using a shared dictionary with model‑specific encoder‑decoder pairs. It preserves activation magnitudes, normalizes only selection scores, and supports single‑model inference via model dropout. Trained on four 1B‑scale language models, SharedSAE retains 96.6% of the mean explained variance of dedicated SAEs, shows higher cross‑model latent correlations, and allows efficient adaptation of new models to the shared latent space.
arXiv:2602. 10352v2 Announce Type: replace-cross Abstract: Self-interpretation methods prompt language models to describe their own internal states, but remain unreliable due to hyperparameter sensitivity.