arXiv:2508.20766v2 Announce Type: replace-cross
Abstract: Safety alignment in Large Language Models (LLMs) often involves mediating internal representations to refuse harmful requests. Recent researc...
By Harethah Abu Shairah, Hasan Abed Al Kader Hammoud, George Turkiyyah, Bernard Ghanem
The paper introduces Latent Space Refusal Anchoring (LSR‑Anchoring), a training‑free technique that extracts a refusal direction from English prompts and applies it to the residual stream of instruction‑tuned models at inference time. The primary variant, Mean‑Activation Steering (MAS), works across several architectures (Llama‑3‑8B, Llama‑3.1‑70B, Mistral‑7B‑Instruct, Qwen2.5‑7B), restoring safety for low‑resource African languages with minimal performance loss, while a refined SAE‑Derived Steering (SDS) further reduces KL divergence without degrading legitimate prompt performance. The method shows positive transfer for Yoruba, Igbo, Igala, and Hausa, but fails for Arabic, suggesting a geometric mismatch rather than a data scarcity issue.
By Godwin Abuh Faruna
arXiv:2511. 11041v2 Announce Type: replace-cross Abstract: We find that current sentence-embedding models produce outputs with a consistent bias: every embedding $e$ decomposes as $\tilde e + \mu$, where the mean $\mu$ is near-identical across all sentences.
By Xingyu Ren, Youran Sun, Haoyu Liang
arXiv:2607. 29539v1 Announce Type: cross Abstract: Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs).
By Gaetano Perrone, Simon Pietro Romano
arXiv:2606. 01196v1 Announce Type: cross Abstract: Safety alignment learned in high-resource languages transfers poorly to low-resource languages.
By Rashad Aziz, Ikhlasul Akmal Hanif, Fajri Koto
arXiv:2607. 14111v1 Announce Type: cross Abstract: Can small language models detect and report on perturbations their own internal activations?
By Ely Hahami, Ishaan Sinha, Lavik Jain