CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits
arXiv:2608. 05732v1 Announce Type: new Abstract: Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment.
arXiv:2607. 07128v1 Announce Type: new Abstract: Language models perform a wide range of tasks at varying levels of abstraction with the capacity to flexibly infer tasks from context, execute multiple tasks simultaneously, and select among competing tasks.
arXiv:2608. 05732v1 Announce Type: new Abstract: Controlling the behavior of large language models (LLMs) remains a critical challenge for AI alignment.
MetaSteer is a new method for steering large language models that learns nonlinear, context-dependent interventions applied to attention projection matrices. Unlike traditional linear, context-independent techniques, MetaSteer adapts its effects based on the input, requiring no linear concept-geometry assumption. Trained once on a pooled preference corpus, it transfers zero‑shot to unseen concepts and out‑of‑distribution contexts, matching or surpassing strong task‑specific baselines on multiple benchmarks and model families.
arXiv:2605. 28854v2 Announce Type: replace-cross Abstract: Large language models (LLMs) exhibit remarkable flexibility in adapting to novel tasks from in-context examples without parameter updates, a capability known as in-context learning (ICL).
arXiv:2606. 15092v1 Announce Type: new Abstract: Activation steering has emerged as a key methodology for controlling the behavior of large language models (LLMs).
arXiv:2607. 02460v1 Announce Type: cross Abstract: Post-training large language models (LLMs) without real-world interaction feedback or human-labeled supervision remains challenging, particularly in specialized domains where expert annotations are costly to obtain.
arXiv:2601. 22594v2 Announce Type: replace-cross Abstract: The high-level concepts that a neural network uses to perform computation need not be aligned to individual neurons (Smolensky, 1986).
arXiv:2606. 08454v1 Announce Type: new Abstract: Activation steering provides a lightweight inference-time mechanism for controlling large language models (LLMs) by modifying their internal activation vectors toward desired behaviors.
Post-training large language models (LLMs) without real-world interaction feedback or human-labeled supervision remains challenging, particularly in specialized domains where expert annotations are costly to obtain. Recent annotation-free self-evolution methods address this by using the model's own outputs as supervision signals, constructing a teacher via additional context and aggregating predictions across multiple rollouts through majority voting to produce pseudo-labels.
arXiv:2606. 05165v1 Announce Type: new Abstract: Training Data Attribution (TDA) seeks to trace a model's predictions back to its training data.
arXiv:2610.01054v1 Announce Type: cross Abstract: In-context learning (ICL) enables language models to perform new tasks from demonstrations without weight updates. However, every ICL inference requi...
arXiv:2608.21664v1 Announce Type: new Abstract: Safe deployment of increasingly capable models will likely come to rely on latent-space monitoring as a complement to behavioral evaluations, especiall...
Fine‑tuning reshapes internal representations of large language models, affecting attention patterns and layer‑wise activations. The study shows that components identified by EAP as important for task performance cluster in specific layers, yet these layers do not align with those undergoing the largest representational changes. Additionally, overlapping EAP components across different tasks do not guarantee cross‑task transfer and can even degrade performance when tasks differ in nature.