arXiv AI By Hadas Orgad, Boyi Wei, Kaden Zheng, Martin Wattenberg, Peter Henderson, Seraphina Goldfarb-Tarrant, Yonatan Belinkov

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types

Read the original on arXiv AI →

arXiv:2604. 09544v2 Announce Type: replace-cross Abstract: Large language models (LLMs) undergo alignment training to avoid harmful behaviors, yet the resulting safeguards remain brittle: jailbreaks routinely bypass them, and fine-tuning on narrow domains can induce ``emergent misalignment'' that generalizes broadly.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 18

The Role of Fine-grained Harm Signals in LLM Safety

The study investigates how category‑specific harmfulness signals, isolated by removing the shared general harmfulness component, influence large language model (LLM) safety. Using activation steering across 11 risk categories in three instruction‑tuned LLMs, the authors find that the presence of harmfulness in these category residuals varies by category and that the pattern of inducing refusal is even more model‑dependent. Additionally, category residuals were shown to enhance the models’ downstream alignment with the shared general harmfulness representation, indicating that fine‑grained signals play a role beyond the general component.

By Soyeon Park (KAIST), Seogyeong Jeong (KAIST), Sunwoo Kim (KAIST), Alice Oh (KAIST)