The paper investigates what aspects of language model behavior are controlled by activation steering. By introducing Cross‑Encoding Steering Evaluation, the authors show that steering effects often follow the extraction index of answer identifiers rather than the semantic content of the answers, especially at deeper layers. They also find that a low‑rank output‑sensitive component captures most of this effect, and that different datasets (NormBank, MNLI, SC101) exhibit varying preferences for extraction‑index versus semantic‑label following.
By Zhiwei Gao, Shaowen Peng, Shoko Wakamiya, Eiji Aramaki
arXiv:2606. 11198v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) systems inject external knowledge to improve LLM outputs, yet the format of injected content -- distinct from its semantic relevance -- can independently distort the model's attention distribution.
By Yuqi Zhang, Di Zhang
arXiv:2608. 02957v1 Announce Type: new Abstract: Steering vectors (SVs) are widely used to influence the expression of concepts (e.
By Max Torop, Aria Masoomi, Jennifer Dy
arXiv:2507. 18043v2 Announce Type: replace-cross Abstract: Inference-time steering methods offer a lightweight alternative to fine-tuning large language models (LLMs) and vision-language models (VLMs) by modifying internal activations at test time without updating model weights.
By Duy Nguyen, Archiki Prasad, Elias Stengel-Eskin, Mohit Bansal
arXiv:2606. 08682v1 Announce Type: cross Abstract: Activation steering has emerged as a popular inference-time technique for modulating the behavior of large language models (LLMs).
By Qi Cao, Jian Lou, Meiting Liu, Wenjie Feng, Dan Li, See-Kiong Ng, Anh Tuan Luu
arXiv:2607. 04223v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level score that does not indicate which sentence is unsupported, or why.
By Mohamed Aly Bouke
arXiv:2607. 29062v1 Announce Type: new Abstract: Model capabilities have improved in large part due to scaling chain of thought.
By Matthew Nguyen, Kyle Cox, Austin Meek, Iv\'an Arcuschin
The paper demonstrates that a single internal direction in modern language models—called the valence axis (V-axis)—captures how positive or negative a sentence feels. By using only nine emotion category names and 50 short narrative paragraphs per emotion, the authors identify this axis via principal component analysis of frozen encoder embeddings, achieving 93% of supervised performance on SST‑2 and strong correlations with human valence ratings across images, audio, and brain recordings. The method transfers across modalities without target‑modality labels, but works only for continuous attributes and is specific to certain model families.
By Yousef Radwan
arXiv:2605. 26795v2 Announce Type: replace Abstract: Chain-of-thought (CoT) prompting enhances large language model performance, yet what drives these gains remains unclear.
By Xiang Wang, Wei Wei
arXiv:2607. 01002v1 Announce Type: cross Abstract: In long-context use, large language models frequently synthesize answers from the meaning of a relevant context span rather than literally copy-pasting them.
By Aryo Pradipta Gema, Beatrice Alex, Pasquale Minervini
arXiv:2408.11827v2 Announce Type: replace
Abstract: Understanding how language models compose meaning from linguistic input remains a central problem in interpretability research. Mechanistic studies...
By Nura Aljaafari, Danilo S. Carvalho, Andr\'e Freitas
arXiv:2608. 14320v1 Announce Type: new Abstract: The anchoring effect is a cognitive bias in which an initial reference value shifts a later judgment toward itself.
By Yiderigun Borjigin, Alexander Hermann, Christian Cyron, Roland Aydin