arXiv AI

TRACE: Trajectory Aware Reasoning for Multi-Turn Adversarial Conversation Evaluation

arXiv:2608. 15594v1 Announce Type: new Abstract: Multi-turn jailbreak attacks have emerged as a critical safety threat to LLMs, as harmful objectives are decomposed across a sequence of apparently benign turns to bypass guardrails.

arXiv AI
Sep 7

AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks

AlcaTRAz is a prompt‑level defense that uses rule trees to insert controlled character‑level perturbations into input text, disrupting jailbreak attacks without modifying or retraining the target LLM. It operates solely on the input, making it suitable for black‑box deployments, and was evaluated on 33 open‑weight models and 22 jailbreak types, outperforming three baseline defenses in 73.4 % of model‑attack combinations. While it significantly reduces high‑severity jailbreak success, it does not eliminate it and is intended as one layer of a broader defense strategy.

By Jakub Re\v{s}, Petr Ka\v{s}ka, Martin Pere\v{s}\'ini, Martin Ukrop, Kamil Malinka
arXiv AI
Aug 6

Temporal Context Awareness: A Defense Framework Against Multi-turn Manipulation Attacks on Large Language Models

arXiv:2503. 15560v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly vulnerable to sophisticated multi-turn manipulation attacks, where adversaries strategically build context through seemingly benign conversational turns to circumvent safety measures and elicit harmful or unauthorized responses.

By Prashant Kulkarni, Assaf Namer
arXiv AI
Aug 26

ADVERSA: Measuring Multi-Turn Guardrail Degradation and Judge Reliability in Large Language Models

The paper introduces ADVERSA, an automated red‑teaming framework that evaluates large language model safety over multiple turns by tracking continuous compliance trajectories instead of binary jailbreak outcomes. Using a fine‑tuned 70B attacker model and a structured 5‑point rubric, the authors conduct controlled experiments on three frontier victim models, measuring guardrail degradation and judge reliability through a triple‑judge consensus. Results show a 26.7% jailbreak rate with most breaches occurring early, and the study documents inter‑judge agreement, attacker drift, and attacker refusals as key factors affecting safety assessment.

By Harry Owiredu-Ashley
arXiv AI
Sep 15

Useless but Safe? Benchmarking Utility Recovery with User Intent Clarification in Multi-Turn Conversations

The paper introduces CarryOnBench, an interactive benchmark that tests whether large language models can revise their interpretation of user intent and recover utility while staying safe in multi‑turn conversations. Using 398 harmful‑looking queries with benign intents, the benchmark simulates 5,970 conversations across 14 models, evaluating both intent‑aligned utility and safety with a new metric called Ben‑Util. Results show that models often withhold information due to misinterpretation, but most can recover with clarifications, revealing failure modes such as unsafe and redundant recovery that single‑turn tests miss.

By Mingqian Zheng, Malia Morgan, Liwei Jiang, Carolyn Rose, Maarten Sap
arXiv AI
Sep 24

Beyond Unsafe Detection: Counterfactually Anchored Evidence Attribution for Multi-Turn LLM Safety Failures

The paper introduces a counterfactually anchored evidence attribution approach for multi‑turn large language model safety failures. It presents a new dataset of 1,762 conversations, including adversarial, benign twins, and high‑risk vocabulary variants, and trains a lightweight hierarchical model that accurately predicts safety violations and attributes them to specific user turns and token spans. The model achieves high detection performance (F1 = 0.988) and significantly reduces adversarial confidence when top‑attributed tokens are removed, while maintaining low false‑positive rates on benign conversations.

By Srinivasan Subramanian, Kazi Aminul Islam, Md. Abdullah Al Hafiz Khan