PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 23496v1 Announce Type: new Abstract: Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards.
Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teaming methods empirically identify effective scenarios through observed attack outcomes, but why particular scenarios weaken refusal remains mechanistically unclear.
arXiv:2511. 19517v3 Announce Type: replace-cross Abstract: Multi-turn conversational attacks, which leverage psychological principles like Foot-in-the-Door (FITD), where a small initial request paves the way for a more significant one, to bypass safety alignments, pose a persistent threat to Large Language Models (LLMs).
arXiv:2608. 11624v1 Announce Type: cross Abstract: Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions.
Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to harmful persuasion becomes a core requirement for reliable behavior.
The paper introduces BLUEPRINT, a safety‑evaluation framework that separates a factorized social‑influence strategy space from WORLDVIEWSIM, a cross‑turn situational context module. Using Monte Carlo Tree Search, it optimizes turn‑level combinations of 18 theory‑grounded influence factors across a four‑turn trajectory, achieving near‑ceiling ASR on six frontier models with an average of only 2.46 queries. The study reveals that model‑specific vulnerabilities arise from distinct influence factors and strategy transitions, yet all models share a recovery pathway that shifts toward concrete, executable task framing to escape hard‑refusal states, highlighting the importance of monitoring how dialogue state makes unsafe requests appear actionable.