The paper challenges the assumption that large language models (LLMs) produce deterministic safety responses by examining how random seeds and temperature settings affect refusal decisions. Across four instruction‑tuned models and 876 harmful prompts, 18‑28% of prompts flipped between refusal and compliance depending on sampling configuration, with higher temperatures reducing decision stability. The authors introduce a Safety Stability Index (SSI) and recommend multi‑sample evaluation protocols that account for stochastic variation rather than relying on single‑shot tests.
By Erik Larsen
The paper investigates whether automatic safety judges evaluate the content of a model’s reply or merely its style. By keeping the reply content fixed and adding various style wrappers—such as educational disclaimers, fake reasoning blocks, or token refusals—the authors show that many judges flip their verdicts, indicating that style can influence safety judgments. The study evaluates over 600 jailbreak examples across multiple judges, revealing that some judges are highly susceptible to style-based manipulation while others remain robust.
By Yongxi Zhou, Wenbo Ye, Yuanzhe Liu, Zihan Dong, Junwei Yao
The paper investigates whether automatic safety judges evaluate the content of a model’s reply or merely its style. By adding content‑invariant style wrappers—such as educational disclaims or token refusals—to fixed replies, the authors show that many judges flip their verdicts, revealing exploitable blind spots. Across more than 600 jailbreak examples and eight judges, some judges exhibit high flip rates (e.g., GPT‑4o‑mini 19.9%) while others remain largely stable, and human validation confirms that most flips are judge errors rather than content changes.
The paper investigates whether existing AI safety benchmarks, designed for large language models, are suitable for evaluating small language models (SLMs). By testing five benchmark suites on 26 open‑source SLMs with a unified scoring rubric, the authors find that ambiguous judgments dominate, especially for complex prompts and certain architectures. This ambiguity, linked to factors like lexical density and output perplexity, undermines the reliability of aggregate leaderboards and reveals a confound between model capability and perceived safety.
By Nyamtulla Shaik, Fengjun Li, Bo Luo
arXiv:2608. 02617v1 Announce Type: cross Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform collecting blinded pairwise preferences alongside multi-criterion rubric ratings.
By Fay Elhassan, David Sasu, Alexandra Kulinkina, Lars Henning Klein, Mary-Anne Hartley
arXiv:2608. 05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks.
By Joshua Fonseca Rivera (Independent), Neil Shah (Independent), David Demitri Africa (UK AI Security Institute), Konstantinos Voudouris (UK AI Security Institute)
arXiv:2609.01210v1 Announce Type: cross
Abstract: Safety benchmarks for large language models often assess the risk of a user query, although the outcome of question answering depends on whether the...
By Rui Yang, Shuang Huang, Junhua Liu, Ziqi Zhao, Qingzhong Yan, Yuhang Sun, Cong Liu, Guoping Hu, Rui Mei, Jing Shao
arXiv:2606. 25487v1 Announce Type: cross Abstract: Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model prompted to grade.
By Yang Gao (Veyon Solutions)
arXiv:2606. 03650v1 Announce Type: cross Abstract: Choosing or ranking language models for a specific application is hardest when no task-specific labeled data exists, and standard public benchmarks cannot be trusted, their items having likely leaked into pretraining, so scores reflect memorization rather than fitness.
By Alexander Apartsin, Yehudit Aperstein
Almost every paper on LLM jailbreaks and prompt injection reports an attack-success rate (ASR), and that number is assigned not by people but by an automated judge: either a safety classifier trained for the task, or a general chat model prompted to grade. The judge is rarely checked.
arXiv:2607. 01153v3 Announce Type: replace-cross Abstract: Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model followed an instruction, refused appropriately, complied with a policy, or misreported progress in an agentic task.
By Brett Reynolds
arXiv:2608. 19266v1 Announce Type: cross Abstract: The OWASP Top 10 for LLM Applications ranks the risks that a community of security practitioners judges most important.
By Kyriakos "Rock" Lambros, Steve Wilson