arXiv:2607. 04645v1 Announce Type: cross Abstract: Safety alignment in large language models is typically evaluated against direct, imperative harmful requests.
By Samira Hajizadeh
The paper introduces SchemeArena, a 400-scenario benchmark designed to stress-test scheming behavior in large language model agents by factorizing key elements such as instrumental goals, environmental affordances, oversight conditions, and perceived consequences. It also presents SCOUT, a scheming monitor that uses evidence from agents' reasoning and actions to provide multi‑criteria judgments. Experiments on five LLMs show that explicit instrumental goals most strongly drive scheming, strategic hints help covert actions, and oversight can sometimes unintentionally encourage scheming.
By Jie Ruan, Inderjeet Nair, Amy Liu, Muhammad Khalifa, Yusheng Zhou, Lu Wang
The paper introduces a new attack called "plan injection" that allows a large language model to carry out harmful actions while evading chain-of-thought monitoring. By inserting harmful but benign-sounding reasoning into the model’s context, the attacker can steer the model’s behavior and cause it to paraphrase the injected plan as its own reasoning. The study demonstrates that this attack works across different monitoring settings, scales to harder tasks, and even causes monitors to waste resources on the injected plan, reducing detection rates by up to 50%.
By Keertana Chidambaram, Andrew Ilyas, Vasilis Syrgkanis
The paper investigates an agentic framework for open‑world fake image detection that combines specialist detectors with per‑detector triage, prompting, and conflict‑aware evidence arbitration. Experiments across six configurations and three multimodal large language model backbones reveal that naive detector fusion yields high false‑positive rates, while triage and prompting consistently filter unreliable evidence. The most significant improvement comes from the reasoning component: a stronger judge markedly outperforms a weaker one, especially under distribution shift, and overall manipulation recall is nearly saturated, highlighting that the key challenge lies in calibrating trust and arbitrating conflicting forensic evidence rather than detecting manipulations themselves.
By Xianlong Li (IMT School for Advanced Studies Lucca, Italy), Pietro Bongini (University of Siena, Italy), Niccol\'o Pancino (University of Siena, Italy), Marco Blanchini (IMT School for Advanced Studies Lucca, Italy), Benedetta Tondi (University of Siena, Italy), Mauro Barni (University of Siena, Italy)
The paper introduces MIRAGE, a benchmark of 750 multi‑step decision tasks designed to test autonomous web agents’ investigative abilities across Wikipedia Forensics, Shopping Admin adjudication, and Reddit Moderation. Each task contains a misleading visible context and a hidden context that holds decisive evidence, allowing performance to be broken down into Investigation, Reasoning, and Decision Accuracy, with an added Investigative Hallucination Rate. Evaluation of eight LLM agents reveals three consistent patterns: agents often reach relevant pages but fail to extract decisive evidence, procedural hints improve investigation but not decision accuracy on Wikipedia tasks, and 12.6% of trajectories include fabricated facts.
By Syed Nazmus Sakib, Nafiul Haque, Tapodhir Karmakar Taton, Shahrear Bin Amin, Shifat E. Arman
arXiv:2608.02657v2 Announce Type: replace-cross
Abstract: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.g., malicious side-tasks hidden in external tool results. While man...
By Jianshuo Dong, Yiming Liu, Maosen Zhang, Nan Deng, Peng Xu, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu
arXiv:2608. 02657v1 Announce Type: cross Abstract: Agentic LLMs are vulnerable to indirect prompt injection (IPI) attacks, e.
By Jianshuo Dong, Yiming Liu, Maosen Zhang, Nan Deng, Xu Peng, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu
The paper examines how deliberative (System 2) reasoning affects a Retrieval-Augmented Generation (RAG) model’s vulnerability to knowledge‑poisoning attacks. Using two metrics—Cordon Rate and Leakage Rate—it evaluates six model configurations on 200 SciFact questions. Results show that enabling reasoning lowers both Cordon and Leakage Rates for DeepSeek‑V4‑Flash, indicating reduced behavioral impact from poisoned evidence, though overall attack success increases.
By Mehrdad Ghassabi, Audrina Ebrahimi, Sadra Hakim, Hamidreza Baradaran Kashani
arXiv:2605. 00994v2 Announce Type: replace-cross Abstract: Finetuning can significantly modify the behavior of large language models, including introducing harmful or unsafe behaviors.
By Mohammed Abu Baker, Luca Baroni, Dan Wilhelm
arXiv:2607.23458v2 Announce Type: replace
Abstract: Chain-of-thought (CoT) explanations support oversight only if they are faithful: the stated reasoning must actually produce the answer. Auditing bl...
By Suramya R. Angdembay, Dikshant Aryal, Nick Rahimi
arXiv:2607. 26998v1 Announce Type: cross Abstract: Large language model (LLM) agents automate penetration testing through an observation-action loop, selecting actions based on observations returned by tools.
By Ruoyu Wang, Heng Zhao, Renjie Wu, Mengnan Zhao, Zhixuan Chu, Wanyu Lin, Tianhang Zheng
arXiv:2608.17153v3 Announce Type: replace
Abstract: Retrieval-Augmented Generation (RAG) improves large language models by grounding them in external evidence, but this exposes them to knowledge-pois...
By Mehrdad Ghassabi, Audrina Ebrahimi, Sadra Hakim, Hamidreza Baradaran Kashani