arXiv:2609.25602v1 Announce Type: new
Abstract: In language models, the choice between believing the prompt and believing the weights is made by a handful of identifiable attention heads. Instruction...
By Shubham Santosh Pandere, Gautam Ranka, Ritika Varshney, Navya Deshmukh, Roushni Sareen, Roshan Kumar Singh
The paper introduces a three-level evaluation framework—behavioral deployment, LM-head readout, and probe recoverability—to distinguish whether a language model fails a syntactic test by not encoding structure or by failing to use it. Using a trilingual control-dependency benchmark, the authors find that probe recoverability consistently exceeds LM-head readout, which in turn exceeds behavioral deployment across seven models and three languages, with the largest gap observed in Qwen3-0.6B Instruct. Layer-localized activation patching shows that instruction tuning shifts the decoded layer later, suggesting decoding favors surface shortcuts and that behavioral evaluation understates what models encode while probing alone overstates what they deploy.
By Zhenyan Lu, He Wang, Xiaohui Huang
arXiv:2609.38205v1 Announce Type: new
Abstract: System prompts are the primary lever practitioners use to control language model behavior, yet what they actually do to the computation inside the tran...
By Muhammad Usama, Dong Eui Chang
The paper investigates where the ‘refusal’ behavior of language models resides across different architectures. It finds that a single direction in the residual stream governs refusal in transformers, and that the same direction—after a rigid rotation—also governs refusal in state‑space models (SSMs). By aligning these directions and applying a detector‑triggered gate, the authors demonstrate that refusal can be effectively transferred across transformer, SSM, recurrent, and hybrid architectures, showing that safety tooling can be ported by re‑estimating the direction at each architecture’s write site rather than rebuilding it from scratch.
By Preethi Carmel Bosco, Gopalakrishnan Srinivasan
arXiv:2606. 22686v2 Announce Type: replace-cross Abstract: Modern Large Language Models (LLMs) rely on extensive safety alignment, yet the mechanistic basis of refusal remains opaque.
By Shivam Ratnakar, Kartikeya Vats
arXiv:2608. 11583v1 Announce Type: new Abstract: Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters.
By Mingyu Zong, Sampad Mohanty, Bhaskar Krishnamachari
arXiv:2608.29109v1 Announce Type: new
Abstract: Large language models often answer structurally unanswerable questions, such as computing cot(-540{\deg}) or evaluating (1).startswith("1"), instead of...
By Yucheng Du, Xiyang Hu
arXiv:2609.00051v1 Announce Type: cross
Abstract: Despite extensive alignment efforts, Large Language Models (LLMs) remain vulnerable to generating unsafe content under adversarial prompting, yet the...
By Kuan-Lin Chu, Chung-En Sun, Tsui-Wei Weng
arXiv:2606. 05378v1 Announce Type: new Abstract: We test whether a single screen-and-ablate recipe -- identify attention-head circuits by task-pattern selectivity, then verify by causal ablation against a matched-random null -- produces consistent mechanistic claims across model families.
By Yongzhong Xu
arXiv:2603. 22473v2 Announce Type: replace-cross Abstract: Hybrid language models combine softmax attention with linear-time sequence mechanisms such as state-space or linear-attention layers, but the functional contribution of each component type remains insufficiently characterized.
By Hector Borobia, Elies Segu\'i-Mas, Guillermina Tormo-Carb\'o
arXiv:2609.00064v1 Announce Type: cross
Abstract: In-context learning (ICL) lets large language models adapt to new tasks from demonstrations, and fine-tuning can erode this behaviour. Many preservat...
By Jinyuan Zhang, Peng He, He Hu, Yin Yuan, ShengShuo Jiao
arXiv:2606.22676v2 Announce Type: replace
Abstract: Refusal on a safety benchmark does not reveal how stable that behavior will remain after model updates. Benign downstream fine-tuning can weaken re...
By Dongyub Jude Lee, Jungseob Lee, Seungyoon Lee, Seongtae Hong, Suhyune Son, Sugyeong Eo, Jaehyung Seo, Heuiseok Lim