The paper introduces the SAST-IR framework to evaluate large language models’ robustness against persuasion attacks in a memory‑less setting, revealing a flaw called "Refusal Inertia" that masks true vulnerability. Using the CP‑Agent and a custom CounterFact‑Strict dataset, the authors demonstrate that simple, diverse attack strategies achieve a 96% success rate, while complex attacks often trigger defensive compliance. The study highlights severe brittleness in current state‑of‑the‑art models when deprived of conversation history.
By Zhuoang Cai
arXiv:2606. 17229v1 Announce Type: cross Abstract: A model that lies while knowing the truth is the central case ELK cannot handle with behavioral evaluation alone.
By Petr Nyoma
The paper introduces GUARD, a method for natural forgetting in large reasoning models that transforms unsafe disclosures into safe-exit trajectories using guided answer‑reasoning distillation. It aligns a frozen model with guidance tokens and distills this behavior into the parameters, aiming for a coherent, non‑disclosing chain of thought followed by a refusal‑style answer. The authors also propose the Natural Forgetting Reasoning Score (NFRS) to evaluate structural stability, fluency, and unsupported substitutes, and demonstrate GUARD’s effectiveness on R‑TOFU and a STAR‑1‑derived harmful‑intent setting.
By Zeyu Yan, Guanghao Zhou, Minghui Qiu, Ming Gao, Cen Chen
A language model's memory can be worse than having no memory at all. Give a model a memory that kept a wrong conclusion but dropped the work behind it, and it emits that stale value as a confident answer; give the same model an empty memory and it abstains.
The paper introduces ‘Fool’s Gold’, a defensive deception technique for open‑weight language models that hardens them against safety‑removal attacks. By training decoy responses within a differentiable simulation of the attack, the method poisons the payoff of stripped refusal mechanisms, producing confident but falsified answers to hazardous requests while preserving benign behavior. Experiments on seven models (9B‑122B) show that 51‑90% of attacked‑state responses become decoys, with the defense accounting for 27‑84% of this effect, and that the defended 122B model remains within benign‑behavior budgets.
By Mark Russinovich
arXiv:2606. 25449v1 Announce Type: cross Abstract: A language model's memory can be worse than having no memory at all.
By Alex Kwon