The study investigates how safety alignment in large language models, trained mainly in English, transfers to other languages. While models show near-perfect harmfulness detection (AUROC > 0.98) using unrelated harmless prompts (easy negatives), performance drops sharply in low‑resource languages when using surface‑similar benign prompts (hard negatives). This degradation persists across multiple languages and models, indicating that easy‑negative evaluation alone cannot confirm cross‑lingual harmfulness representation quality.
By Paras Balani, Subhrakanta Panda
arXiv:2606. 01196v1 Announce Type: cross Abstract: Safety alignment learned in high-resource languages transfers poorly to low-resource languages.
By Rashad Aziz, Ikhlasul Akmal Hanif, Fajri Koto
arXiv:2609.25602v1 Announce Type: new
Abstract: In language models, the choice between believing the prompt and believing the weights is made by a handful of identifiable attention heads. Instruction...
By Shubham Santosh Pandere, Gautam Ranka, Ritika Varshney, Navya Deshmukh, Roushni Sareen, Roshan Kumar Singh
The paper investigates where the ‘refusal’ behavior of language models resides across different architectures. It finds that a single direction in the residual stream governs refusal in transformers, and that the same direction—after a rigid rotation—also governs refusal in state‑space models (SSMs). By aligning these directions and applying a detector‑triggered gate, the authors demonstrate that refusal can be effectively transferred across transformer, SSM, recurrent, and hybrid architectures, showing that safety tooling can be ported by re‑estimating the direction at each architecture’s write site rather than rebuilding it from scratch.
By Preethi Carmel Bosco, Gopalakrishnan Srinivasan
arXiv:2609.14861v1 Announce Type: cross
Abstract: A deployed language model may refuse a harmful request in English yet comply with its faithful translation, revealing a cross-lingual safety failure...
By Mohammed Ahnouch, Lotfi Elaachack
arXiv:2609.14754v1 Announce Type: cross
Abstract: Causal claims about large language model (LLM) internals rest on measurements. Those might include a projection, a cosine, an ablation delta, or an i...
By Orion Reblitz-Richardson