arXiv Computation and Language

The Geometry of Harmfulness in Multi-Turn Attacks

arXiv Computation and Language
Sep 16

Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion

The paper introduces the SAST-IR framework to evaluate large language models’ robustness against persuasion attacks in a memory‑less setting, revealing a flaw called "Refusal Inertia" that masks true vulnerability. Using the CP‑Agent and a custom CounterFact‑Strict dataset, the authors demonstrate that simple, diverse attack strategies achieve a 96% success rate, while complex attacks often trigger defensive compliance. The study highlights severe brittleness in current state‑of‑the‑art models when deprived of conversation history.

By Zhuoang Cai
arXiv Computation and Language
3d ago

Evaluating Language Model Safety Across Long Adversarial Conversations

The paper investigates how conversational safety in language models degrades over extended, adversarial interactions. By testing three instruction‑tuned models with persistent adversarial users across up to 101 turns, the study finds that safe‑response rates drop sharply from 85–100% at the first turn to 15–44% by the end. This demonstrates that strong single‑turn safety does not guarantee continued safety in long conversations.

By Parisa Salmani, Peter R. Lewis
arXiv AI
Jul 7

Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types

arXiv:2604. 09544v2 Announce Type: replace-cross Abstract: Large language models (LLMs) undergo alignment training to avoid harmful behaviors, yet the resulting safeguards remain brittle: jailbreaks routinely bypass them, and fine-tuning on narrow domains can induce ``emergent misalignment'' that generalizes broadly.

By Hadas Orgad, Boyi Wei, Kaden Zheng, Martin Wattenberg, Peter Henderson, Seraphina Goldfarb-Tarrant, Yonatan Belinkov
Hugging Face Trending Papers
Jul 2

Safety Targeted Embedding Exploit via Refinement

Safety training for large language models (LLMs) is conducted predominantly in English, leaving uncertain how well safety mechanisms generalize to low-resource languages and mixed-language code-switching. We show that this creates an epistemic gap in which models confidently generate harmful responses for inputs that fall outside the distribution of their safety training.