Hugging Face Blog

Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

arXiv Computation and Language
Sep 7

Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal

The paper introduces a boundary-aware self‑distillation framework for controlled large language model safety refusal, addressing the need for different refusal boundaries within the same topic. It combines controlled topic generation, coverage repair, in‑distribution compensation data, and harmful‑benign pairs to train and evaluate refusal behavior. Experiments on Qwen3‑8B show that escalating retries dramatically improve target‑domain refusal rates while reducing unsafe responses, though they also increase over‑refusal, highlighting the trade‑off between safety and usability.

By Alejo L\'opez-\'Avila, Iker Garc\'ia-Ferrero, Jezabel Garcia, Antonio Tiene, Rom\'an Or\'us
arXiv Computation and Language
Sep 24

SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems

SafeTutors is a benchmark designed to evaluate both safety and pedagogical effectiveness of AI tutoring systems across mathematics, physics, and chemistry. It introduces a risk taxonomy of 11 harm dimensions and 48 sub‑risks based on learning‑science literature, focusing on issues such as answer over‑disclosure, misconception reinforcement, and loss of scaffolding. The study finds that all tested models exhibit broad harms, that larger scale does not mitigate these issues, and that multi‑turn interactions significantly increase pedagogical failures from 17.7% to 77.8%.

By Rima Hazra, Bikram Ghuku, Ilona Marchenko, Yaroslava Tokarieva, Sayan Layek, Somnath Banerjee, Julia Stoyanovich, Mykola Pechenizkiy
arXiv AI
Sep 7

Refuse without Refusal: A Structural Analysis of Safety-Tuning Responses for Reducing False Refusals in Language Models

The paper investigates how large language models balance helpfulness and safety by refusing harmful queries while responding to benign ones. It decomposes safety-tuning responses into a boilerplate refusal statement and a rationale, finding that the statement causes false refusals by relying on superficial cues. Training on rationales alone reduces false refusals without compromising safety performance, suggesting that fine‑grained safety supervision is essential for better alignment.

By Minji Kim, Hyounghun Kim