arXiv:2606. 00152v1 Announce Type: cross Abstract: LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users.
By Mingxuan Zhang, Jiahui Han, Dadi Guo, Songze Li, Guanchu Wang, Na Zou, Dongrui Liu, Xia Hu
arXiv:2606. 28061v1 Announce Type: cross Abstract: Large language models (LLMs) have increasingly moved from standalone text generation systems to agents that invoke external tools, access environments, and execute multi-step tasks.
By Shijing Hu, Liang Liu, Zhu Meng, Zhicheng Zhao
CIPL (Channel Inversion for Privacy Leakage) is a channel-aware framework designed to evaluate black-box privacy leakage in large language model agents. It models the leakage process through stages of sensitive source, selection, assembly, execution, observation, and extraction, assessing how selected sensitive units become attacker-recoverable outputs. Experiments across memory, retrieval, and tool-mediated targets, plus a live-agent case study, reveal that recoverability depends on factors beyond storage labels, such as observation surface, prompt alignment, retrieval depth, and provider behavior, and that a semantic audit can uncover disclosures missed by exact matching.
By Tao Huang, Guosen Wu, Guolong Zheng, Jiayang Meng, Chen Hou, Xu Yang, Xuechao Yang, Feng Xia
arXiv:2607. 19449v1 Announce Type: cross Abstract: Evaluation frameworks for tool-augmented LLM agents focus overwhelmingly on capability metrics or explicit tool crashes, leaving silent infrastructure failures and HTTP 200 responses with empty, null, or malformed payloads largely unaudited.
By Aarushi Singh
The paper introduces SARA, a framework that separates action induction from runtime authorization in tool‑augmented LLM agents. By treating these as distinct roles, SARA uses an Action Probe to record action provenance and only authorizes tool calls that align with the user objective and past successful executions. Experiments on AgentDojo and AgentDyn show that SARA reduces action‑to‑side‑effect risk to below 0.63% while preserving task performance.
By Xiaokun Guo, Zhen Xu, Dongdong Huo, Yanqiu Zhang, Wei Wang, Qinfu Yang, Dongjin Yu, Yu Wang
arXiv:2512. 16310v3 Announce Type: replace-cross Abstract: LLM-based agents increasingly use multiple external tools to complete complex tasks.
By Yuxuan Qiao, Dongqin Liu, Hongchang Yang, Wei Zhou, Songlin Hu