arXiv AI

Activation Steering Induces Emergent Misalignment: A More Comprehensive Evaluation

arXiv:2606. 08682v1 Announce Type: cross Abstract: Activation steering has emerged as a popular inference-time technique for modulating the behavior of large language models (LLMs).

arXiv AI
Sep 21

Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models

Fine‑tuning reshapes internal representations of large language models, affecting attention patterns and layer‑wise activations. The study shows that components identified by EAP as important for task performance cluster in specific layers, yet these layers do not align with those undergoing the largest representational changes. Additionally, overlapping EAP components across different tasks do not guarantee cross‑task transfer and can even degrade performance when tasks differ in nature.

By Lingfang Li, Procheta Sen, Shubham Das, Danushka Bollegala