The paper investigates whether existing AI safety benchmarks, designed for large language models, are suitable for evaluating small language models (SLMs). By testing five benchmark suites on 26 open‑source SLMs with a unified scoring rubric, the authors find that ambiguous judgments dominate, especially for complex prompts and certain architectures. This ambiguity, linked to factors like lexical density and output perplexity, undermines the reliability of aggregate leaderboards and reveals a confound between model capability and perceived safety.
By Nyamtulla Shaik, Fengjun Li, Bo Luo
arXiv:2608.20554v1 Announce Type: cross
Abstract: The critical failure modes in deployed large language models (LLMs) are cross-dimensional: a model can score 99.3 in safety alignment while refusing...
By Fatih Deniz, Yazan Boshmaf, Dorde Popovic, Issa Khalil
arXiv:2606. 09401v1 Announce Type: new Abstract: Recent work has applied differential privacy (DP) to adapt large language models (LLMs) for sensitive applications, offering theoretical guarantees.
By Bart{\l}omiej Marek, Lorenzo Rossi, Vincent Hanke, Xun Wang, Michael Backes, Franziska Boenisch, Adam Dziedzic
arXiv:2501. 14940v4 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with human values is essential for their safe deployment and widespread adoption.
By Guangzhi Sun, Xiao Zhan, Shutong Feng, Philip C. Woodland, Jose Such
The paper examines how the scores of cybersecurity large language model (LLM) benchmarks vary depending on the evaluation pipeline used. By auditing eight benchmarks across ten different LLMs, the authors uncover 15 systematic failure modes and demonstrate that a single pipeline choice can shift a model’s score by over 80 percentage points, significantly altering rankings. They also show that even semantically similar tasks can produce different model rankings due to incompatible evaluation conventions, and that standardizing pipelines can move most models by at least three ranks on at least one benchmark.
By Aymene Berriche, Cathrine Shalby, Mohannad Alhanahnah, Yazan Boshmaf
arXiv:2607. 28801v1 Announce Type: cross Abstract: Benchmark datasets are central to evaluating Large Language Models (LLMs), yet they are typically conceived as monolithic tasks, obscuring substantial variation in the demands of individual samples.
By Philipp D. Siedler, Jordan Sassoon