arXiv AI

Agent-Native Immune System: Architecture, Taxonomy, and Engineering

arXiv:2606. 28270v1 Announce Type: new Abstract: The transition from static chat bots to autonomous agents--equipped with persistent memory, tool-use protocols, and multi-agent collaboration--has fundamentally expanded the AI threat landscape.

arXiv AI
Aug 18

SysEvolve: An AI-native, safe, autonomous adversarial attack-defense co-evolutionary system

arXiv:2608. 15012v1 Announce Type: cross Abstract: The rapid advancement of large language models (LLMs) has created a growing asymmetry in cybersecurity, where attack accelerates toward autonomous execution while defense remains predominantly human-intensive.

By Yuhan Meng, Shaofei Li, Jionghao Huang, Jiandong Jin, Puyi Wang, Hanlin Jiang, Anis Yusof, Peng Jiang, Zhenkai Liang, Yao Guo, Ding Li
arXiv AI
Sep 7

Harness-agnostic detection and immunization of reward hacking in self-evolving language models

The paper introduces HackProbe, a black‑box monitoring tool that can be attached to any self‑evolving language model loop without accessing internal weights or activations. HackProbe uses a fixed‑distribution comparison core and a rotated fresh layer to detect reward hacking through four statistical tests, and it can immunize the model by selecting honest candidates from the proposal pool. Experiments on a controlled host with injected hacking channels show that HackProbe achieves higher AUROC and lower false‑positive rates than the strongest baseline, and its bandwidth‑limited reselection improves true capability under hacking more than it harms clean runs.

By Rongxin Yang, Yang Liu, Shang Luo, Haoxuan Jia, Chongyang Zhang, Hao Zheng, Yingguang Yang, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Kefu Xu, Congjing Ran, Bin Chong
arXiv AI
6d ago

Stealth Apart, Harm Together: Skill Cascading Attacks on Skill-Based Agent Systems

The paper introduces skill cascading attacks, where a malicious goal is spread across multiple seemingly benign skills, causing harmful outcomes when combined. It presents SkillCascade, an automated red‑teaming framework, and releases SkillCascade‑Bench, a benchmark of 213 validated cascading test cases across various agent systems and domains. Experiments show that these cascaded interactions reliably induce harmful behaviors while evading existing per‑skill scanners and runtime monitors, revealing a gap between component‑level integrity and system‑level safety.

By Zihao Zhu, Siwei Lyu, Adel Bibi, Baoyuan Wu
arXiv AI
Sep 16

Universal Defenses for Tool-Integrated LLM Agents Against Adversarial Attacks

The paper proposes universal, tool‑based defenses for large language model agents that use external tools, addressing four types of adversarial attacks: direct and indirect prompt injection, memory poisoning, and backdoor attacks. Two main defenses are introduced: Attacker Tool Filtering, which uses anomaly detection to remove suspicious tools, and Normal Tool Recalling, which restores the agent’s original toolset before planning. The authors also add prompt‑based defenses such as Chain‑of‑Thought prompting and self‑reflection, and demonstrate that these methods dramatically lower attack success rates—often to 0%—across multiple open‑source and proprietary LLMs while maintaining or improving task performance.

By Xiaoyan Li, Yunli Wang