arXiv AI By Mingyue Cui, Linghui Shen, Xingyi Yang

SAE Interventions are Unreliable: Post-Intervention Recovery of Suppressed Behavior

Read the original on arXiv AI →

arXiv:2606. 18322v1 Announce Type: cross Abstract: Sparse Autoencoders (SAEs) decompose residual-stream activations into interpretable features.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 4

ObserverBench: Testing Mechanistic Estimates for Intervention and Control

ObserverBench is a benchmark framework that evaluates whether internal mechanistic estimators—called observers—are suitable for guiding interventions, control, or safety actions in language models. It separates estimation accuracy from the loss incurred by the chosen action, showing that accurate predictions do not always lead to better decisions. Experiments on GPT‑2‑small, Qwen2.5‑7B, Gemma‑2‑9B‑it, and Qwen3.5‑9B demonstrate that observers trained on action loss can reduce deployment loss, while traditional metrics like AUROC may rank monitors differently from actual performance.

By Vijay Erramilli
arXiv AI
6d ago

FLIP: Final Layer Inference-Time Probing for Vision-Language Models

FLIP is a final‑layer inference‑time probe designed to test whether a logit‑facing intervention site in an open‑weight vision‑language model (VLM) supports structured, task‑linked computation rather than generic perturbation. The probe applies elementwise flooring to the final normalized hidden state before logit computation, leaving other model components unchanged. By sweeping intervention strength on a controlled detection/counting task, FLIP identifies three regimes—negligible change, a bounded interior regime with improved detection recall and reduced counting error, and over‑suppression—while a four‑criterion protocol ensures the observed effects are mechanistically interpretable.

By Drandreb Earl O. Juanico, Rowel O. Atienza