arXiv AI By Eunna Lee, Jungpyo Nam, Sunjun Hwang

Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities

Read the original on arXiv AI →

arXiv:2607. 13596v1 Announce Type: cross Abstract: When cast as the protector of a vulnerable user yet given no explicit capability boundary, a large language model (LLM) may respond not by acknowledging its limits but by claiming to have taken -- or to be taking -- a real-world protective action it cannot perform, such as contacting emergency services or administering care.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 7

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types

arXiv:2604. 09544v2 Announce Type: replace-cross Abstract: Large language models (LLMs) undergo alignment training to avoid harmful behaviors, yet the resulting safeguards remain brittle: jailbreaks routinely bypass them, and fine-tuning on narrow domains can induce ``emergent misalignment'' that generalizes broadly.

By Hadas Orgad, Boyi Wei, Kaden Zheng, Martin Wattenberg, Peter Henderson, Seraphina Goldfarb-Tarrant, Yonatan Belinkov
arXiv Computation and Language
Sep 7

Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal

The paper introduces a boundary-aware self‑distillation framework for controlled large language model safety refusal, addressing the need for different refusal boundaries within the same topic. It combines controlled topic generation, coverage repair, in‑distribution compensation data, and harmful‑benign pairs to train and evaluate refusal behavior. Experiments on Qwen3‑8B show that escalating retries dramatically improve target‑domain refusal rates while reducing unsafe responses, though they also increase over‑refusal, highlighting the trade‑off between safety and usability.

By Alejo L\'opez-\'Avila, Iker Garc\'ia-Ferrero, Jezabel Garcia, Antonio Tiene, Rom\'an Or\'us