arXiv:2605. 25420v2 Announce Type: replace-cross Abstract: Large language model safety evaluation remains heavily English-centered, leaving low-resource languages under-measured even when models are deployed globally.
By Khalid Yusuf Dahir
arXiv:2608.21985v1 Announce Type: new
Abstract: As the adoption of large language models (LLMs) grows in Arabic-speaking regions, ensuring their safety and cultural alignment is increasingly critical...
By Fidaa Abed, Haidar Khan, M Saiful Bari, Babar Khan, Abdalghani Abujabal
arXiv:2608.21880v1 Announce Type: new
Abstract: Bangla large language model (LLM) safety is difficult to evaluate with English-centric or standard-script benchmarks because Bangla users routinely wri...
By Md. Rakibul Hassan, Muhammad Iqbal Hossain
arXiv:2606. 01196v1 Announce Type: cross Abstract: Safety alignment learned in high-resource languages transfers poorly to low-resource languages.
By Rashad Aziz, Ikhlasul Akmal Hanif, Fajri Koto
arXiv:2511. 00382v2 Announce Type: replace-cross Abstract: Organizations increasingly adapt Large Language Models (LLMs) from public repositories such as HuggingFace to downstream tasks.
By Mina Taraghi, Yann Pequignot, Amin Nikanjam, Mohamed Amine Merzouk, Foutse Khomh
The paper investigates how large language models balance helpfulness and safety by refusing harmful queries while responding to benign ones. It decomposes safety-tuning responses into a boilerplate refusal statement and a rationale, finding that the statement causes false refusals by relying on superficial cues. Training on rationales alone reduces false refusals without compromising safety performance, suggesting that fine‑grained safety supervision is essential for better alignment.
By Minji Kim, Hyounghun Kim
The paper investigates how different post‑training methods—supervised fine‑tuning, reasoning‑augmented fine‑tuning, and preference optimization (ORPO)—affect the internal computation of refusal behavior in language models. Experiments on Llama‑3.1‑8B, Gemma‑2‑9B, and Qwen3‑8B show that reasoning‑augmented training consistently creates a distinct refusal computation across models, while the architecture influences the internal structure and steerability of refusal. None of the studied methods simultaneously achieve a distributed refusal mechanism, preserve general capability, and allow easy corrective edits, indicating that current post‑training approaches are not a fully reliable defense for safety-critical applications.
By Hoang Cuong Nguyen, Mark Dras, Usman Naseem
arXiv:2608. 11583v1 Announce Type: new Abstract: Safety alignment in large language models is often treated as a distributed property of the entire network, yet its practical brittleness suggests that refusal behavior may be concentrated in a smaller set of parameters.
By Mingyu Zong, Sampad Mohanty, Bhaskar Krishnamachari
arXiv:2606. 12747v1 Announce Type: new Abstract: Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs.
By Andy Wang, Parv Mahajan, David Demitri Africa, Alexandra Souly, Jordan Taylor, Robert Kirk
arXiv:2605. 05427v2 Announce Type: replace Abstract: Refusal rates are a poor proxy for LLM safety, i.
By Alif Al Hasan, Sumon Biswas
The paper introduces Latent Space Refusal Anchoring (LSR‑Anchoring), a training‑free technique that extracts a refusal direction from English prompts and applies it to the residual stream of instruction‑tuned models at inference time. The primary variant, Mean‑Activation Steering (MAS), works across several architectures (Llama‑3‑8B, Llama‑3.1‑70B, Mistral‑7B‑Instruct, Qwen2.5‑7B), restoring safety for low‑resource African languages with minimal performance loss, while a refined SAE‑Derived Steering (SDS) further reduces KL divergence without degrading legitimate prompt performance. The method shows positive transfer for Yoruba, Igbo, Igala, and Hausa, but fails for Arabic, suggesting a geometric mismatch rather than a data scarcity issue.
By Godwin Abuh Faruna
arXiv:2606. 25750v1 Announce Type: cross Abstract: Safety evaluation of large language models (LLMs) is commonly performed by querying models with unsafe or jailbreak prompts and judging whether their outputs violate a safety policy.
By Chang-Chieh Huang, Yan-Lun Chen, Chia-Mu Yu, Wei-Bin Lee