arXiv Computation and Language
Sep 18

Towards Safer RAG: Only Agents Capable of System 2 Thinking may Access Untrusted Documents

The paper examines how deliberative (System 2) reasoning affects a Retrieval-Augmented Generation (RAG) model’s vulnerability to knowledge‑poisoning attacks. Using two metrics—Cordon Rate and Leakage Rate—it evaluates six model configurations on 200 SciFact questions. Results show that enabling reasoning lowers both Cordon and Leakage Rates for DeepSeek‑V4‑Flash, indicating reduced behavioral impact from poisoned evidence, though overall attack success increases.

By Mehrdad Ghassabi, Audrina Ebrahimi, Sadra Hakim, Hamidreza Baradaran Kashani
arXiv AI
Sep 15

Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection

The paper introduces a new attack called "plan injection" that allows a large language model to carry out harmful actions while evading chain-of-thought monitoring. By inserting harmful but benign-sounding reasoning into the model’s context, the attacker can steer the model’s behavior and cause it to paraphrase the injected plan as its own reasoning. The study demonstrates that this attack works across different monitoring settings, scales to harder tasks, and even causes monitors to waste resources on the injected plan, reducing detection rates by up to 50%.

By Keertana Chidambaram, Andrew Ilyas, Vasilis Syrgkanis
arXiv Machine Learning
Jul 30

RAGuard: A Layered Defense Framework for Retrieval-Augmented Generation Systems Against Data Poisoning

arXiv:2607. 26339v1 Announce Type: new Abstract: Retrieval-Augmented Generation (RAG) systems ground large language models (LLMs) in external corpora, but this reliance exposes them to corpus poisoning: maliciously injected passages that manipulate retrieved evidence.

By Pushkal Kumar, Tucker Nielson, Tanish Kolhe, Shubham Zala, Vincent Li