arXiv:2608.30041v1 Announce Type: cross
Abstract: Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later pri...
By Wujie Xiong, Rabimba Karanjai, Yang Lu, Weidong Shi, Lei Xu
arXiv:2609.05770v1 Announce Type: new
Abstract: Differentially private (DP) fine-tuning methods treat sparse Mixture-of-Experts (MoE) models as a single dense block, ignoring that shared layers see a...
By Duc Dm, Khai Le-Duc, Nguyen Do, Minh Son Hoang, Florent Draye, Thai Hoang, Hoang Phuong Dam, Jiarui Liu, Chris Ngo, Terry Jingchen Zhang, Anh Le Duc Tran, Nhat Do Minh, Minh Ngoc Le, My T. Thai, Ran Xu, Silvio Savarese, Mona Diab, Bernhard Sch\"olkopf, Zhijing Jin, Huy L. Nguyen, Daeyoung Kim
arXiv:2604. 16870v2 Announce Type: replace-cross Abstract: AI agents increasingly call external tools (file system, network, APIs) through the Model Context Protocol (MCP).
By Daeyeon Son
arXiv:2606. 09549v1 Announce Type: cross Abstract: Tool-using large language model (LLM) agents face two distinct security failures: unauthorized external actions and exposure of sensitive plaintext inside the runtime before any final output check can intervene.
By Yuhan Ma, Stefan Schmid
arXiv:2609.36570v1 Announce Type: cross
Abstract: Indirect prompt injection makes an LLM agent treat untrusted retrieved text as instructions. We present CounterSteer, an inference-time defense that...
By Mark Russinovich
arXiv:2606. 26057v1 Announce Type: cross Abstract: AI agents are granted access to tools, APIs, and other infrastructure, making them active principals in those systems.
By Seth Dobrin, {\L}ukasz Chmiel
arXiv:2608.23224v1 Announce Type: cross
Abstract: Retrieval can efficiently and effectively augment a frozen vision--language--action (VLA) policy without retraining, yet retrieved text becomes a con...
By Zhiruo Zhou, Zelin Li, Xiwen Chen, Jiazhuo Li, Chenwei Wang, Huiming Chen, Xiaojun Zhu
arXiv:2606. 27567v1 Announce Type: cross Abstract: Prompt injection is the top security risk for LLM-integrated applications, yet every defense proposed so far has been broken.
By Dewank Pant, Shruti Lohani, Avijit Kumar
ServeGuard is a supply‑chain primitive that allows a publisher to ship a proof‑carrying adapter for an open‑weight language model, proving in zero‑knowledge that the adapter contains no hidden backdoor channel in the monitor’s blind subspace. The proof is inexpensive because it relies on a deterministic function of the public base model, and the served residual is the model’s own public floor. The system lets consumers or regulators verify the absence of this class of hidden channels without revealing the certified read factor or trusting the publisher.
By Dominik Dahlem, Rui Vieira
The paper reports a privacy breach in a two-node split‑LLM training system where the returned gradient reveals which data rows were real, despite the system passing standard privacy checks. By exploiting the fact that decoy rows produce zero gradients, an attacker can identify real rows with 100% accuracy across multiple runs. The authors demonstrate that adding gradient clipping and noise can mitigate the leak, but the system remains vulnerable to several untested attack vectors.
By Georgios Politis, Evangelos Pappas
arXiv:2609.22243v1 Announce Type: new
Abstract: Input-conditioned neural interventions raise a runtime question: what persists when one behavioral specification admits multiple actions whose validity...
By Xianliang Zeng, Zhanzhan Zhao
arXiv:2609.06036v1 Announce Type: new
Abstract: Proposal-based controllers---learned policies, language-model planners, and other black-box \emph{generators}---are increasingly deployed behind runtim...
By Guangxi Wan, Yongbo Xie, Yuqi Liu, Qingwei Dong, Qingxin Li, Hongfei Bai, Peng Zeng