arXiv AI By Trilok Padhi, Pinxian Lu, Abdulkadir Erol, Tanmay Sutar, Gauri Sharma, Mina Sonmez, Munmun De Choudhury, Ugur Kursuncu

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks

Read the original on arXiv AI →

arXiv:2510. 14207v3 Announce Type: replace Abstract: Large Language Model (LLM) agents are powering a growing share of interactive web applications, yet remain vulnerable to misuse and harm.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

TIER: Threat Implicitness Benchmark for Evaluating LLM Safety Behaviors

The paper introduces TIER, a Threat Implicitness Benchmark designed to evaluate large language model (LLM) safety behaviors across four risk domains and four threat levels, ranging from explicit harmful requests to sophisticated jailbreaks. Responses are scored on a six-label behavior scale by two independent LLM judges. Experiments on six open-weight LLMs reveal that safety behaviors change gradually with threat level, contextual prompts produce the most varied responses, and jailbreaks expose significant robustness gaps, underscoring the importance of behavior-aware safety evaluation.

By Thu-Hien Trinh-Thi, Hai-Yen Vong, Thanh-Ha Ung-Dung, Tram Ho
arXiv AI
Sep 3

Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking

The paper introduces BLUEPRINT, a safety‑evaluation framework that separates a factorized social‑influence strategy space from WORLDVIEWSIM, a cross‑turn situational context module. Using Monte Carlo Tree Search, it optimizes turn‑level combinations of 18 theory‑grounded influence factors across a four‑turn trajectory, achieving near‑ceiling ASR on six frontier models with an average of only 2.46 queries. The study reveals that model‑specific vulnerabilities arise from distinct influence factors and strategy transitions, yet all models share a recovery pathway that shifts toward concrete, executable task framing to escape hard‑refusal states, highlighting the importance of monitoring how dialogue state makes unsafe requests appear actionable.

By Siyu Chen, Haoran Wang, Xiaojian Li, Yao Huang, Yinpeng Dong, Wei Xu