arXiv AI By Hadas Orgad, Boyi Wei, Kaden Zheng, Martin Wattenberg, Peter Henderson, Seraphina Goldfarb-Tarrant, Yonatan Belinkov

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types

Read the original on arXiv AI →

arXiv:2604. 09544v2 Announce Type: replace-cross Abstract: Large language models (LLMs) undergo alignment training to avoid harmful behaviors, yet the resulting safeguards remain brittle: jailbreaks routinely bypass them, and fine-tuning on narrow domains can induce ``emergent misalignment'' that generalizes broadly.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.