arXiv AI By Yuhang Wang

PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection

Read the original on arXiv AI →

arXiv:2607. 16199v1 Announce Type: new Abstract: Multi-agent LLM systems increasingly rely on a Planner to decompose goals into sub-task sequences that downstream Executor and Critic agents execute and audit.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 16

Universal Defenses for Tool-Integrated LLM Agents Against Adversarial Attacks

The paper proposes universal, tool‑based defenses for large language model agents that use external tools, addressing four types of adversarial attacks: direct and indirect prompt injection, memory poisoning, and backdoor attacks. Two main defenses are introduced: Attacker Tool Filtering, which uses anomaly detection to remove suspicious tools, and Normal Tool Recalling, which restores the agent’s original toolset before planning. The authors also add prompt‑based defenses such as Chain‑of‑Thought prompting and self‑reflection, and demonstrate that these methods dramatically lower attack success rates—often to 0%—across multiple open‑source and proprietary LLMs while maintaining or improving task performance.

By Xiaoyan Li, Yunli Wang
arXiv AI
Sep 25

Understanding and Exploiting Initialization Anchoring Weakness in Feedback-Based Agent Planning

The paper investigates a vulnerability in feedback‑based agent planning, showing that the first round of feedback corrects a large portion of adversarial directions (46%) while subsequent rounds see a sharp decline (13% and 7%). The authors attribute this to an initialization anchoring weakness driven by plausible plan shifts, lack of counterevidence, and persistence of accepted directions. They introduce “InitAnchor”, a black‑box attack framework that exploits these factors, achieving high attack success rates across diverse tasks, architectures, and LLMs, and remaining effective against multiple defenses and real‑world agents.

By Chuanchao Zang, Jianing Wang, Wenyu Chen, Xiangtao Meng, Li Wang, Xinyu Gao, Peng Zhan, Zheng Li, Shanqing Guo
arXiv AI
Aug 20

Breaking Planner Integrity Boundary: Enviroment State-Text Injection Attack on LLM-Driven Embodied Agents

The paper introduces the Environment State-Text Injection (ESTI) attack, a novel method that manipulates the textual representation of environment states in large language model‑driven embodied agents without altering user instructions, model parameters, or executors. ESTI re‑frames adversarial goals as false state evidence that aligns with the current environment, thereby influencing both planning and execution through object properties, spatial relations, affordances, task‑stage rules, and execution feedback. The authors also present ESTI‑Bench, a benchmark that evaluates attack propagation across the planning‑to‑execution closed loop, and demonstrate that ESTI outperforms existing baselines on multiple embodied task datasets, achieving up to 89.32% higher planning‑level and 43.69% higher execution‑level attack success rates.

By Jiawei Liu, Jiacheng Guo, Tian Zhang, Yiwei Xu, Juan Wang, Jinlin Fan, Bowen Xiao, Chi Guo, Keyan Guo, Hongxin Hu