The study investigates how safety alignment in large language models, trained mainly in English, transfers to other languages. While models show near-perfect harmfulness detection (AUROC > 0.98) using unrelated harmless prompts (easy negatives), performance drops sharply in low‑resource languages when using surface‑similar benign prompts (hard negatives). This degradation persists across multiple languages and models, indicating that easy‑negative evaluation alone cannot confirm cross‑lingual harmfulness representation quality.
By Paras Balani, Subhrakanta Panda
arXiv:2606. 01196v1 Announce Type: cross Abstract: Safety alignment learned in high-resource languages transfers poorly to low-resource languages.
By Rashad Aziz, Ikhlasul Akmal Hanif, Fajri Koto
arXiv:2609.25602v1 Announce Type: new
Abstract: In language models, the choice between believing the prompt and believing the weights is made by a handful of identifiable attention heads. Instruction...
By Shubham Santosh Pandere, Gautam Ranka, Ritika Varshney, Navya Deshmukh, Roushni Sareen, Roshan Kumar Singh
The paper investigates where the ‘refusal’ behavior of language models resides across different architectures. It finds that a single direction in the residual stream governs refusal in transformers, and that the same direction—after a rigid rotation—also governs refusal in state‑space models (SSMs). By aligning these directions and applying a detector‑triggered gate, the authors demonstrate that refusal can be effectively transferred across transformer, SSM, recurrent, and hybrid architectures, showing that safety tooling can be ported by re‑estimating the direction at each architecture’s write site rather than rebuilding it from scratch.
By Preethi Carmel Bosco, Gopalakrishnan Srinivasan
arXiv:2609.14861v1 Announce Type: cross
Abstract: A deployed language model may refuse a harmful request in English yet comply with its faithful translation, revealing a cross-lingual safety failure...
By Mohammed Ahnouch, Lotfi Elaachack
arXiv:2609.14754v1 Announce Type: cross
Abstract: Causal claims about large language model (LLM) internals rest on measurements. Those might include a projection, a cosine, an ablation delta, or an i...
By Orion Reblitz-Richardson
The paper evaluates how three large mixture‑of‑experts models (Alibaba, OpenAI, NVIDIA) can be fine‑tuned to reason in a low‑resource language, specifically Greek. Accuracy metrics show little change, but the authors uncover significant qualitative improvements: after supervised fine‑tuning, models reason in Greek on ~98% of items, with better grammaticality and retained general ability. Reinforcement learning with pre‑registered rewards further eliminates reasoning‑channel leaks and format skips, while the Greek‑reasoning habit remains robust to an accuracy‑only gradient.
By Ayoub Kirouane, Christos Petrocheilos
arXiv:2602. 06941v2 Announce Type: replace-cross Abstract: Large language models can recover mid-generation from task-misaligned activation steering, producing explicit verbal restarts (e.
By Alex McKenzie, Keenan Pepper, Stijn Servaes, Martin Leitgab, Murat Cubuktepe, Mike Vaiana, Diogo de Lucena, Judd Rosenblatt, Michael S. A. Graziano
When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions.
arXiv:2608. 09095v1 Announce Type: new Abstract: Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial intelligence.
By Shuyi Miao, Wangjie Qiu, Pengyang Shao, Canran Xiao, Fei Shen, Zhiming Zheng, Tat-Seng Chua
The paper introduces Latent Space Refusal Anchoring (LSR‑Anchoring), a training‑free technique that extracts a refusal direction from English prompts and applies it to the residual stream of instruction‑tuned models at inference time. The primary variant, Mean‑Activation Steering (MAS), works across several architectures (Llama‑3‑8B, Llama‑3.1‑70B, Mistral‑7B‑Instruct, Qwen2.5‑7B), restoring safety for low‑resource African languages with minimal performance loss, while a refined SAE‑Derived Steering (SDS) further reduces KL divergence without degrading legitimate prompt performance. The method shows positive transfer for Yoruba, Igbo, Igala, and Hausa, but fails for Arabic, suggesting a geometric mismatch rather than a data scarcity issue.
By Godwin Abuh Faruna
arXiv:2609.00378v1 Announce Type: new
Abstract: Large language models pay a well-documented tax on non-English text: the same content costs several times more tokens, and because attention is quadrat...
By Madhulatha Mandarapu, Sandeep Kunkunuru