arXiv AI By Yang Liu, Bin Chong, Wenkai Yang, Shuai Zhang, Yancheng Chen, Feiyu Han, GuoZhen, Cheng Zhang, Huaibing Xie, Changze Lv, Shihan Dou, Pluto Zhou

Certified Multi-Turn Robustness for LLM Safety via Compositional Bounds and Safety Persistence

Read the original on arXiv AI →

arXiv:2608. 20820v1 Announce Type: new Abstract: Large language models (LLMs) are vulnerable to multi-turn jailbreak attacks that progressively manipulate conversation context.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
3d ago

Evaluating Language Model Safety Across Long Adversarial Conversations

The paper investigates how conversational safety in language models degrades over extended, adversarial interactions. By testing three instruction‑tuned models with persistent adversarial users across up to 101 turns, the study finds that safe‑response rates drop sharply from 85–100% at the first turn to 15–44% by the end. This demonstrates that strong single‑turn safety does not guarantee continued safety in long conversations.

By Parisa Salmani, Peter R. Lewis