The paper investigates whether existing AI safety benchmarks, designed for large language models, are suitable for evaluating small language models (SLMs). By testing five benchmark suites on 26 open‑source SLMs with a unified scoring rubric, the authors find that ambiguous judgments dominate, especially for complex prompts and certain architectures. This ambiguity, linked to factors like lexical density and output perplexity, undermines the reliability of aggregate leaderboards and reveals a confound between model capability and perceived safety.
By Nyamtulla Shaik, Fengjun Li, Bo Luo
Safety benchmarks for large language models often assess the risk of a user query, although the outcome of question answering depends on whether the response violates a policy. This distinction is cri...
arXiv:2609.01210v1 Announce Type: cross
Abstract: Safety benchmarks for large language models often assess the risk of a user query, although the outcome of question answering depends on whether the...
By Rui Yang, Shuang Huang, Junhua Liu, Ziqi Zhao, Qingzhong Yan, Yuhang Sun, Cong Liu, Guoping Hu, Rui Mei, Jing Shao
The paper investigates how large language models balance helpfulness and safety by refusing harmful queries while responding to benign ones. It decomposes safety-tuning responses into a boilerplate refusal statement and a rationale, finding that the statement causes false refusals by relying on superficial cues. Training on rationales alone reduces false refusals without compromising safety performance, suggesting that fine‑grained safety supervision is essential for better alignment.
By Minji Kim, Hyounghun Kim
arXiv:2605. 17173v2 Announce Type: replace-cross Abstract: Large language models exhibit safety degradation in non-English languages.
By Max Zhang, Ameen Patel, Sang T. Truong, Sanmi Koyejo
arXiv:2606. 25476v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated remarkable performance across natural language processing tasks, yet their deployment in high-stakes applications raises critical concerns regarding reliability, safety, and trustworthiness.
By Abrar Alotaibi, Raed Mughus, Moataz Ahmed
arXiv:2606. 28843v1 Announce Type: cross Abstract: Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task.
By Will Hawkins, Kaivalya Rawal, Jonathan Rystr{\o}m, Stratis Tsirtsis, Zihao Fu, Greta Warren, Ryan Brown, Eoin Delaney, Sandra Wachter, Brent Mittelstadt, Chris Russell
arXiv:2608.21775v1 Announce Type: new
Abstract: Large Language Models (LLMs) are increasingly deployed in real-world applications, yet they remain vulnerable to generating harmful content. From adver...
By Afshin Orojlooyjadid, Hitesh Patel
Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content. In this work, we reveal that this advanced safety awareness inadvertently introduces a fatal vulnerability.
WildSEEK is a new dataset of 3,000 real user information‑seeking queries, manually annotated for risk‑sensitive domains and whether the query is factoid or analytical. The accompanying evaluation framework tests LLM responses against four failure criteria—sycophantic behavior, overreliance, a default US‑centric perspective, and poor handling of vulnerable populations—finding higher failure rates for analytical queries. The authors also train classifiers on WildSEEK to analyze over 1.8 million realistic queries, revealing that more than a third are high‑risk and often analytical.
By Tanise Ceron, Joachim Baumann, Elisa Bassignana, Berat Cabuk, Dirk Hovy, Debora Nozza
arXiv:2608. 13695v1 Announce Type: cross Abstract: Large language model providers routinely cite multilingual safety benchmarks spanning a dozen or more languages as evidence that their models are safe for non-English-speaking users.
By Chialuka Prisca-Mary Onuoha, Bright Etornam Sunu, Rashidat Sikiru
RAG-Safety-Bench is a benchmark designed to evaluate how retrieval-augmented generation (RAG) affects the safety of large language models (LLMs). It isolates safety impacts by testing four conditions: non-RAG, RAG with an oracle document, RAG with related but non-answer documents, and RAG with random safe documents. Results on five open-source LLMs reveal an inverse relationship between benign and unsafe capabilities, show that baseline safety guardrails do not guarantee safety in RAG, and confirm that even benign documents can trigger unsafe generation.
By Adithiyan Rajan Indira Saravanan, Kathleen C. Fraser