arXiv Machine Learning

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits

arXiv:2608. 05732v1 Announce Type: new Abstract: Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment.

arXiv Machine Learning
Jul 9

Distributed Sparse Interventions in Language Models

arXiv:2607. 07128v1 Announce Type: new Abstract: Language models perform a wide range of tasks at varying levels of abstraction with the capacity to flexibly infer tasks from context, execute multiple tasks simultaneously, and select among competing tasks.

By Maximilian S. Ernst (Max Planck School of Cognition, Center for Lifespan Psychology Max Planck Institute for Human Development, Machine Learning Group Technische Universit\"at Berlin), Lorenz Linhardt (Machine Learning Group Technische Universit\"at Berlin, Berlin Institute for the Foundations of Learning and Data), Aaron Peikert (Center for Lifespan Psychology Max Planck Institute for Human Development), Oliver Eberle (Machine Learning Group Technische Universit\"at Berlin, Berlin Institute for the Foundations of Learning and Data)
arXiv Machine Learning
Jun 2

Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time

arXiv:2606. 01923v1 Announce Type: cross Abstract: Large Language Models (LLMs) frequently exhibit "contextual disregard" when faced with input evidence that conflicts with their internal parametric memory, leading to persistent factual hallucinations.

By Mingkuan Zhao, Yide Gao, Wentao Hu, Suquan Chen, Tianchen Huang, Zhenhua An, Zetao Chang, Xiayu Sun, Yuheng Min
arXiv AI
Sep 2

What Drives Representation Steering? A Mechanistic Case Study on Steering Refusal

The paper investigates how steering vectors influence large language models (LLMs) by conducting a mechanistic case study on refusal behavior. Using a multi-token activation patching framework, the authors find that steering methods primarily target the OV circuit of the attention mechanism, largely ignoring the QK circuit, and that these circuits are functionally interchangeable across different steering approaches. The study also shows that steering vectors can be sparsified by 85–96% with minimal performance loss and that key dimensions are consistently identified across methods.

By Stephen Cheng, Sarah Wiegreffe, Dinesh Manocha