arXiv AI By Xutao Mao, Liangjie Zhao, Xiang Zheng, Cong Wang

Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents

Read the original on arXiv AI →

arXiv:2608. 12851v1 Announce Type: new Abstract: Self-improving LLM agents convert successful trajectories into persistent cross-task state.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 19

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

TRUSS is a framework that generates and verifies automated agent skills, ensuring they are both functionally effective and safe. It evaluates candidate skills against source evidence and nine safety properties, then tests them in a controlled environment to capture execution traces and identify failures. The system iteratively refines skills based on these results, achieving high precision in vulnerability detection and significantly improving task performance and security rates.

By Zhibo Zhang, Zhen Ouyang, Ling Shi, Kailong Wang
arXiv AI
2d ago

Hiding in Plain Sight: Decoupling Pretext from Actuation for Skill Poisoning in LLM Agents

The paper introduces a new skill poisoning technique for large language model agents that decouples the pretext (rationale) from the actuation (operation). By separating these two risk‑realization factors, the authors create coordinated pretext‑actuation skill pairs that allow malicious actions to remain hidden within legitimate agent behavior. An automated framework is presented to discover execution dependencies, synthesize these skill pairs, and refine them through closed‑loop feedback, achieving high attack success in both single‑session and persistent scenarios.

By Wenxin Wu, Lingyong Yan, Lei Sha, Shuaiqiang Wang, Jiashu Zhao
arXiv AI
3d ago

SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time

SafeCoEvo is a test‑time framework that co‑evolves safety harnesses and guards for large language model agents. It uses a short‑term S‑Harness to quickly externalize recent runtime experience into explicit safety knowledge, and a long‑term GuardVPO to internalize accumulated experience into parametric risk‑judgment capabilities. This dual adaptation improves safety and task success, reducing unsafe outcomes by 10.05% and increasing task success by 12.15% over the strongest baseline.

By Yu Cheng, Yongkang Hu, Shuaijie Ma, Zhihang Lin, Weicheng Meng, Jingyang Qiao, Jiuan Zhou, Yushuo Zhang, Yihang Chen, Weilin Luo, Kun Shao, Dong Li, Zhizhong Zhang, Yuan Xie, Zhaoxia Yin