arXiv:2508.15526v2 Announce Type: replace
Abstract: The rapid proliferation of large language models (LLMs) has intensified the requirement for reliable safety evaluation to uncover model vulnerabili...
By Xiangyang Zhu, Yuan Tian, Chunyi Li, Kaiwei Zhang, Wei Sun, Guangtao Zhai
PASTABench introduces a benchmark of 1,139 multi-turn trajectories to evaluate proactive safety monitoring in large language models. It formalizes three dimensions of intervention—whether, when, and what risk—to address gaps in step-level isolation and post-hoc trajectory assessment. The study finds that proactive intervention is largely unsolved, with the best model achieving only 40.74% optimal-timing interventions, and reveals that smaller models’ safety scores are often driven by lexical overfitting rather than true risk comprehension.
By Jiapeng Sun, Yujin Zhou, Han Zhu, Pengcheng Wen, Jiayi Zhou, Sirui Han, Yike Guo
arXiv:2508. 00923v3 Announce Type: replace Abstract: Large language models (LLMs) are increasingly used to answer health-related questions and support healthcare workflows, yet evidence for their safety still relies heavily on static benchmarks that can rapidly become obsolete or be optimized against.
By Jiazhen Pan (Cherise), Bailiang Jian (Cherise), Paul Hager (Cherise), Yundi Zhang (Cherise), Che Liu (Cherise), Friederike Jungmann (Cherise), Hongwei Bran Li (Cherise), Julian Canisius (Cherise), Chenyu You (Cherise), Junde Wu (Cherise), Jiayuan Zhu (Cherise), Fenglin Liu (Cherise), Yuyuan Liu (Cherise), Niklas Bubeck (Cherise), Moritz Knolle (Cherise), Chen (Cherise), Chen (Cherise), Christian Wachinger, Zhenyu Gong, Cheng Ouyang, Georgios Kaissis, Benedikt Wiestler, Daniel Rueckert
arXiv:2603. 07445v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often require fine-tuning (FT) to perform well on downstream tasks, but FT can induce safety-alignment drift even when the training dataset contains only benign data.
By Guoli Wang, Haonan Shi, Tu Ouyang, An Wang
Safety alignment of large language models (LLMs) typically depends on high-quality supervision data, such as safe demonstrations or preference pairs. However, in real-world deployment, emerging safety requirements are often specified as natural-language policies, while corresponding supervision data may be costly, delayed, or unavailable.
arXiv:2606. 00027v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed across healthcare, yet existing benchmarks fail to capture model behavior under adversarial or ethically complex conditions common in clinical practice.
By Andrei Marian Feier, Veysel Kocaman, Yigit Gul, Ahmet Korkmaz, Alexander Thomas, Aleksei Zakharov, Jay Gil, Mehmet Butgul, David Talby