arXiv:2608. 04286v1 Announce Type: cross Abstract: Large language models (LLMs) are often used in conjunction with external knowledge sources to improve their factual accuracy and decrease hallucinations, through methods such as Retrieval-Augmented Generation (RAG).
By Atri Vivek Sharma, Brian Formento, Alessio Lomuscio
arXiv:2605. 21706v2 Announce Type: replace Abstract: Safety-aligned language models are trained to refuse harmful requests, yet refusal behavior can be suppressed by steering their internal representations.
By Giorgio Piras, Raffaele Mura, Fabio Brau, Maura Pintor, Luca Oneto, Fabio Roli, Battista Biggio
arXiv:2503. 06269v3 Announce Type: replace-cross Abstract: Traditional white-box methods for creating adversarial perturbations against LLMs typically rely only on gradient computation from the targeted model, ignoring the internal mechanisms responsible for attack success or failure.
By Thomas Winninger, Boussad Addad, Katarzyna Kapusta
arXiv:2606. 01441v1 Announce Type: new Abstract: Large language models (LLMs) excel in reasoning and knowledge-intensive tasks but remain vulnerable to prompt-level adversarial attacks that preserve intent while triggering commonsense hallucinations.
By Boxuan Wang, Zhuoyun Li, Xiaowei Huang, Yi Dong
arXiv:2510. 02999v5 Announce Type: replace-cross Abstract: Existing gradient-based jailbreak attacks on Large Language Models (LLMs) typically optimize adversarial suffixes to align the LLM output with predefined target responses.
By Xinzhe Huang, Wenjing Hu, Tianhang Zheng, Kedong Xiu, Hongsheng Hu, Xiaojun Jia, Di Wang, Zhan Qin, Kui Ren
arXiv:2607. 28959v1 Announce Type: cross Abstract: Adversarial training is one of the most effective defenses against adversarial attacks, yet the computational cost remains prohibitive at modern scales, especially for large language models (LLMs).
By Weiyi He, Yuping Lin, Jiliang Tang, Yue Xing
arXiv:2607. 19894v1 Announce Type: cross Abstract: Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs.
By Yuxi Li, Zhibo Zhang, Kailong Wang, Xingshuo Han, Ling Shi, Haoyu Wang
arXiv:2609.23596v1 Announce Type: cross
Abstract: With face-recognition models now embedded in everyday authentication and surveillance, recent works have pinpointed a critical weakness: these models...
By Ben Shapira, Roi Cohen, Shang-Tse Chen, Mahmood Sharif
The paper introduces OverThink, a slowdown attack that forces reasoning language models (RLMs) to produce many more reasoning tokens while still giving correct answers. By injecting decoy reasoning problems—such as Markov decision processes, language translation, or graphic comprehension—into the model’s context, attackers can dramatically increase token generation (up to 46× on SQuAD and 17× on coding agents). The study evaluates the attack on both proprietary and open-source RLMs across multiple datasets, explores multimodal and coding‑agent variants, and tests several defenses, concluding that defending against OverThink is challenging and that newer RLMs are even more vulnerable due to higher per‑token costs and increased reasoning token usage.
By Abhinav Kumar, Jaechul Roh, Ali Naseh, Marzena Karpinska, Mohit Iyyer, Amir Houmansadr, Eugene Bagdasarian
The paper introduces the SAST-IR framework to evaluate large language models’ robustness against persuasion attacks in a memory‑less setting, revealing a flaw called "Refusal Inertia" that masks true vulnerability. Using the CP‑Agent and a custom CounterFact‑Strict dataset, the authors demonstrate that simple, diverse attack strategies achieve a 96% success rate, while complex attacks often trigger defensive compliance. The study highlights severe brittleness in current state‑of‑the‑art models when deprived of conversation history.
By Zhuoang Cai
arXiv:2609.22228v1 Announce Type: cross
Abstract: Large language models increasingly rely on long chain-of-thought (CoT) trajectories for complex reasoning, but autoregressive generation brings subst...
By Shaolong Chen, Ang Li, Mingjie Li, Yisen Wang
Large language models (LLMs) are vulnerable to backdoor attacks, where hidden triggers induce malicious outputs. Existing defenses generally fall into inference-time detection or training-time mitigation, but face two key limitations.