arXiv AI By Fengpeng Li, Kemou Li, Qizhou Wang, Haiwei Wu, Jiantao Zhou, Di Wang

TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety

Read the original on arXiv AI →

The paper introduces TRACE, a token‑level objective designed to reduce multi‑turn safety risks in large language models. TRACE assigns each token a weight based on the discounted return of a refusal‑attributable advantage, comparing a frozen reference model with a refusal‑ablated copy to credit early tokens for later refusal evidence. Evaluated across five open‑weight models and seven multi‑turn attacks, TRACE achieves the lowest attack success rate in all 35 model‑attack pairs while maintaining model utility within 1.23 points on MMLU and HellaSwag.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 16

Benchmarking Factual Robustness of LLMs via Multi-conversation Persuasion

The paper introduces the SAST-IR framework to evaluate large language models’ robustness against persuasion attacks in a memory‑less setting, revealing a flaw called "Refusal Inertia" that masks true vulnerability. Using the CP‑Agent and a custom CounterFact‑Strict dataset, the authors demonstrate that simple, diverse attack strategies achieve a 96% success rate, while complex attacks often trigger defensive compliance. The study highlights severe brittleness in current state‑of‑the‑art models when deprived of conversation history.

By Zhuoang Cai
arXiv AI
Sep 24

Beyond Unsafe Detection: Counterfactually Anchored Evidence Attribution for Multi-Turn LLM Safety Failures

The paper introduces a counterfactually anchored evidence attribution approach for multi‑turn large language model safety failures. It presents a new dataset of 1,762 conversations, including adversarial, benign twins, and high‑risk vocabulary variants, and trains a lightweight hierarchical model that accurately predicts safety violations and attributes them to specific user turns and token spans. The model achieves high detection performance (F1 = 0.988) and significantly reduces adversarial confidence when top‑attributed tokens are removed, while maintaining low false‑positive rates on benign conversations.

By Srinivasan Subramanian, Kazi Aminul Islam, Md. Abdullah Al Hafiz Khan
arXiv Machine Learning
1d ago

Refusal Localizes, the Damage Relocates: Safety Layers Under Few-Sample Fine-Tuning

The paper investigates how fine‑tuning large language models with a small number of harmful examples can erode their refusal behavior, and explores whether localizing safety‑related behavior to specific layers or directions can provide robust defenses. Experiments across six checkpoints from four model families show that harmful and benign prompts remain linearly separable after attack, and that patching clean hidden states or freezing layers up to a transition depth can restore refusal. However, attackers can bypass these defenses by spreading updates or targeting singular directions, indicating that adaptive fine‑tuning can defeat localized repairs and highlighting the need for multiple defensive checks.

By Jungseob Lee, Dongyub Jude Lee, Sugyeong Eo, Seongtae Hong, Seungyoon Lee, Heuiseok Lim
arXiv AI
Sep 15

Overflip: Repetition-Induced Label Flips in Guardrail Models

Guardrail models, which screen malicious prompts in LLM services, often use lightweight Transformers with short context windows and bucketed positional encodings. The study identifies a new failure mode called Overflip, where repeating a prompt causes the guardrail’s prediction to flip from malicious to benign as the sequence length increases. Experiments on nine popular guardrails show that 5 models exhibit MAL→BEN flips on 100 prompts, with flip rates ranging from 8% to 92% and first flips occurring between 2.6k and 9.4k tokens, highlighting a gradual attention dispersion distinct from traditional attention‑dilution attacks.

By Xu He, Chih-Hsuan Lin, Hung-Mao Chen, Junjie Xiong, Yan Zhai, Kun Sun