Production LLMs must handle inputs that attempt to override system instructions, bypass safety policies or elicit harmful responses. A common mitigation is a separate guardrail model. Existing reports...
arXiv:2607. 02079v1 Announce Type: cross Abstract: We present HaloGuard 1.
By Navaneeth Sangameswaran, Preetham S, Ashmiya Lenin
arXiv:2608.21880v1 Announce Type: new
Abstract: Bangla large language model (LLM) safety is difficult to evaluate with English-centric or standard-script benchmarks because Bangla users routinely wri...
By Md. Rakibul Hassan, Muhammad Iqbal Hossain
The paper introduces Trustworthy RAG, an evaluation agent designed to detect misinformation and knowledge poisoning in Retrieval-Augmented Generation systems. It combines natural language inference verification, a five-signal poison detector, and a weighted Trust Index to assess the reliability of retrieved content. Experiments on multiple LLMs show high accuracy and precision, with the agent effectively blocking unsafe advice in a secure-coding assistant scenario.
By Balkrishna Giri, Md Toufique Hasan, Jussi Rasku, Muhammad Waseem, Pekka Abrahamsson
arXiv:2608. 08468v1 Announce Type: cross Abstract: Agent Skills---structured packages of instructions and scripts that augment LLM-based agents---are rapidly proliferating, yet their security properties remain under-explored.
By Xinze Chen, Chi Zhang, Ping Ji, Yimin Liu
arXiv:2606. 02240v1 Announce Type: cross Abstract: Indirect prompt injection in tool-use agents is a concrete production threat: LLM agents read from integrations (third-party services such as Gmail, Salesforce, or Jira accessed through tool calls) whose response content the user neither writes nor controls.
By Hiskias Dingeto, Will Leeney
arXiv:2608. 19901v1 Announce Type: cross Abstract: Agent Skills extend LLM agents with reusable instruction packages that may also include scripts, resources, and service configuration.
By Yue Wang, Yi Liu, Gelei Deng, Ying Zhang, Yuekang Li, Zhenyu Chen, Leo Zhang
arXiv:2604. 06550v3 Announce Type: replace-cross Abstract: Agent skills combine natural-language instructions with executable code while inheriting an agent's filesystem, credential, and network access.
By Yinghan Hou, Zongyou Yang
arXiv:2508. 10029v3 Announce Type: replace-cross Abstract: Safety-aligned large language models can still be manipulated through white-box interventions that modify their internal representations.
By Wenpeng Xing, Bohan Yang, Mohan Li, Chunqiang Hu, Haitao Xu, Ningyu Zhang, Bo Lin, Meng Han
arXiv:2604. 07223v2 Announce Type: replace-cross Abstract: As large language models (LLMs) evolve from static chatbots into autonomous agents, the primary vulnerability surface shifts from final outputs to intermediate execution traces.
By Yen-Shan Chen, Sian-Yao Huang, Cheng-Lin Yang, Yun-Nung Chen
arXiv:2608.30041v1 Announce Type: cross
Abstract: Large language model agents place outputs from external skills into their execution context, allowing attacker-controlled data to influence later pri...
By Wujie Xiong, Rabimba Karanjai, Yang Lu, Weidong Shi, Lei Xu
The paper critiques current encoded‑prompt safety benchmarks that focus only on harmful requests, showing that such tests can misrepresent a model’s safety. By evaluating the benign arm under the same encoding, the authors reveal a substantial drop in the harm gap—sometimes to zero—indicating that the encoding masks true safety deficiencies. Across multiple large models and training pipelines, they document that the encoding can either hide or falsely inflate safety metrics, and they identify twelve specific instrument defects that contribute to these misleading results.
By Haoyu Zhang, Haowen Xu, Xiao Luo, Hanwen Liu, Yang Chen, Zijian Xiao, Yi Feng, Xiangchen Guan, Mohammad Zandsalimy, Shanu Sushmita