arXiv:2609.14861v1 Announce Type: cross
Abstract: A deployed language model may refuse a harmful request in English yet comply with its faithful translation, revealing a cross-lingual safety failure...
By Mohammed Ahnouch, Lotfi Elaachack
arXiv:2608. 02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form.
By Yongxi Zhou, Junwei Yao, Yuanzhe Liu, Zihan Dong, Wenbo Ye, Jiaxi Wen, Lai Yun Choi
arXiv:2606. 01196v1 Announce Type: cross Abstract: Safety alignment learned in high-resource languages transfers poorly to low-resource languages.
By Rashad Aziz, Ikhlasul Akmal Hanif, Fajri Koto
arXiv:2605. 17173v2 Announce Type: replace-cross Abstract: Large language models exhibit safety degradation in non-English languages.
By Max Zhang, Ameen Patel, Sang T. Truong, Sanmi Koyejo
Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change. This has a known blind spot, a systematic misreading survives the perturbation and gets certified wrong, which we show is computable, not just real: an error is invisible to an edit exactly when the two commute, so the errors a suite cannot reach form its joint centralizer, a set that shrinks as edits are added and can be written down rather than guessed at.
arXiv:2608. 05675v1 Announce Type: new Abstract: Label-free reliability for vision-language models rests on invariance: perturb the input and a faithful reader's answer should not change.
By Rasul Khanbayov, Hasan Kurban
arXiv:2608. 09095v1 Announce Type: new Abstract: Uncovering the internal mechanisms underlying the safety capabilities of large language models (LLMs) is crucial for developing trustworthy artificial intelligence.
By Shuyi Miao, Wangjie Qiu, Pengyang Shao, Canran Xiao, Fei Shen, Zhiming Zheng, Tat-Seng Chua
arXiv:2608. 03842v1 Announce Type: cross Abstract: When a language model fails on surface-perturbed input (typos, OCR noise, homophones), "which layer is responsible" has three natural operationalizations: where representations diverge most (sensitivity), where restoring clean activations recovers the prediction (causality), and where a small adapter can repair the damage (compensatory capacity) - and we show these three layer maps dissociate.
By Nathan Labiosa, David Buff, Ena Nayak, Erica Donno
The study investigates how safety alignment in large language models, trained mainly in English, transfers to other languages. While models show near-perfect harmfulness detection (AUROC > 0.98) using unrelated harmless prompts (easy negatives), performance drops sharply in low‑resource languages when using surface‑similar benign prompts (hard negatives). This degradation persists across multiple languages and models, indicating that easy‑negative evaluation alone cannot confirm cross‑lingual harmfulness representation quality.
By Paras Balani, Subhrakanta Panda
arXiv:2607. 15218v1 Announce Type: new Abstract: Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world.
By Weimeng Wang, Ziqiang Wang, Zihang Zhan, Chuanpu Fu, Qi Li, Ke Xu
arXiv:2606. 07559v2 Announce Type: replace-cross Abstract: Fine-tuning a language model often fails silently when its correct completion must outrank a near-synonym competitor.
By Vaibhav Prakash, Jayasri Dontabhaktuni
arXiv:2608. 08032v1 Announce Type: new Abstract: Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the same request in a lower-resource language.
By Ramakrishna P. Kompella, Aadit Mahajan