arXiv:2609.06951v1 Announce Type: cross
Abstract: Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activa...
By Srikanth Malla, Chiho Choi, Joon Hee Choi
arXiv:2607. 25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time.
By Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal, Dipesh Mahato
arXiv:2609.39971v1 Announce Type: cross
Abstract: Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action,...
By Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee
The paper investigates how small action errors evolve when using action chunking in behavioural cloning. By injecting errors at each state and observing their growth under open‑loop (no replanning) and closed‑loop (replanning) regimes, the authors classify states as contracting, expanding, or unresolved. Across twelve manipulation tasks, they find that stable states are rare, error amplification is common, and that short‑horizon fitting can overestimate long‑horizon propagation. Predictors trained on camera and proprioceptive data can recover open‑loop stability but only partially capture closed‑loop dynamics, indicating that standard imitation learning does not reliably produce policies that contract errors when perturbed.
By Aryan Goyal
The paper investigates whether fine‑tuning a language model erases previously embedded activation steering interventions that suppress refusals and encourage brevity. Across five instruction‑tuned models (3B–14B) subjected to non‑adversarial supervised fine‑tuning (SFT) and reinforcement learning from human feedback (RLHF), the authors find that the steering’s behavioural effect degrades when the fine‑tuning objective conflicts with the targeted behaviour, yet the underlying weight edits remain largely unchanged. Mechanistically, the steering vectors survive with minimal alteration, but functionally the steering is vulnerable and must be re‑validated after downstream training.
By Philipp E. Glass, Allan Tucker, Yongmin Li, Alina Miron
The paper investigates where the ‘refusal’ behavior of language models resides across different architectures. It finds that a single direction in the residual stream governs refusal in transformers, and that the same direction—after a rigid rotation—also governs refusal in state‑space models (SSMs). By aligning these directions and applying a detector‑triggered gate, the authors demonstrate that refusal can be effectively transferred across transformer, SSM, recurrent, and hybrid architectures, showing that safety tooling can be ported by re‑estimating the direction at each architecture’s write site rather than rebuilding it from scratch.
By Preethi Carmel Bosco, Gopalakrishnan Srinivasan