The paper introduces a boundary-aware self‑distillation framework for controlled large language model safety refusal, addressing the need for different refusal boundaries within the same topic. It combines controlled topic generation, coverage repair, in‑distribution compensation data, and harmful‑benign pairs to train and evaluate refusal behavior. Experiments on Qwen3‑8B show that escalating retries dramatically improve target‑domain refusal rates while reducing unsafe responses, though they also increase over‑refusal, highlighting the trade‑off between safety and usability.
By Alejo L\'opez-\'Avila, Iker Garc\'ia-Ferrero, Jezabel Garcia, Antonio Tiene, Rom\'an Or\'us
How we think about safety for users experiencing mental or emotional distress, the limits of today’s systems, and the work underway to refine them.
arXiv:2605. 05427v2 Announce Type: replace Abstract: Refusal rates are a poor proxy for LLM safety, i.
By Alif Al Hasan, Sumon Biswas
SafeTutors is a benchmark designed to evaluate both safety and pedagogical effectiveness of AI tutoring systems across mathematics, physics, and chemistry. It introduces a risk taxonomy of 11 harm dimensions and 48 sub‑risks based on learning‑science literature, focusing on issues such as answer over‑disclosure, misconception reinforcement, and loss of scaffolding. The study finds that all tested models exhibit broad harms, that larger scale does not mitigate these issues, and that multi‑turn interactions significantly increase pedagogical failures from 17.7% to 77.8%.
By Rima Hazra, Bikram Ghuku, Ilona Marchenko, Yaroslava Tokarieva, Sayan Layek, Somnath Banerjee, Julia Stoyanovich, Mykola Pechenizkiy
The paper investigates how large language models balance helpfulness and safety by refusing harmful queries while responding to benign ones. It decomposes safety-tuning responses into a boilerplate refusal statement and a rationale, finding that the statement causes false refusals by relying on superficial cues. Training on rationales alone reduces false refusals without compromising safety performance, suggesting that fine‑grained safety supervision is essential for better alignment.
By Minji Kim, Hyounghun Kim
Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content. In this work, we reveal that this advanced safety awareness inadvertently introduces a fatal vulnerability.
arXiv:2609.14796v1 Announce Type: new
Abstract: The threat that AI persuasion poses to human control has been acknowledged in the literature, but not yet systematically studied. Now that persuasion a...
By Joshua Levy, Mick Yang, Kellin Pelrine
arXiv:2606. 07874v1 Announce Type: new Abstract: LLMs-as-judges are the only way to evaluate safety at scale.
By Anissa Alloula, Federico Licini, Ava Batchkala, Seraphina Goldfarb-Tarrant
arXiv:2608.21775v1 Announce Type: new
Abstract: Large Language Models (LLMs) are increasingly deployed in real-world applications, yet they remain vulnerable to generating harmful content. From adver...
By Afshin Orojlooyjadid, Hitesh Patel