Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic
Read the original on Hugging Face Blog →The Flow has not summarised this story yet — read it at Hugging Face Blog.
The Flow has not summarised this story yet — read it at Hugging Face Blog.
The paper introduces a boundary-aware self‑distillation framework for controlled large language model safety refusal, addressing the need for different refusal boundaries within the same topic. It combines controlled topic generation, coverage repair, in‑distribution compensation data, and harmful‑benign pairs to train and evaluate refusal behavior. Experiments on Qwen3‑8B show that escalating retries dramatically improve target‑domain refusal rates while reducing unsafe responses, though they also increase over‑refusal, highlighting the trade‑off between safety and usability.
How we think about safety for users experiencing mental or emotional distress, the limits of today’s systems, and the work underway to refine them.
arXiv:2605. 05427v2 Announce Type: replace Abstract: Refusal rates are a poor proxy for LLM safety, i.
SafeTutors is a benchmark designed to evaluate both safety and pedagogical effectiveness of AI tutoring systems across mathematics, physics, and chemistry. It introduces a risk taxonomy of 11 harm dimensions and 48 sub‑risks based on learning‑science literature, focusing on issues such as answer over‑disclosure, misconception reinforcement, and loss of scaffolding. The study finds that all tested models exhibit broad harms, that larger scale does not mitigate these issues, and that multi‑turn interactions significantly increase pedagogical failures from 17.7% to 77.8%.