arXiv AI

Unicode TAG-Block Concealment of Tool-Metadata Payloads in the Model Context Protocol: An Approval-View Fidelity Gap Across Three Independent Server Implementations

arXiv:2607. 05744v1 Announce Type: cross Abstract: The Model Context Protocol (MCP) is the dominant way coding agents discover and invoke external tools.

arXiv AI
Aug 26

Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors

The paper introduces Semantic Overlays, a steering technique that adds non‑textual annotations to a language model’s input by applying learned adapters at specific prefill positions. These overlays create an out‑of‑band channel that encodes span identity and complex semantics, enabling the model to interpret marked text differently—such as rewriting code in a specified language or ignoring executable instructions. Experiments show that Semantic Overlays dramatically reduce prompt‑injection success rates while preserving model utility and readability of marked spans.

By Joshua Penman
arXiv AI
Jun 2

Attested Tool-Server Admission: A Security Extension to the Model Context Protocol

arXiv:2605. 24248v2 Announce Type: replace-cross Abstract: The Model Context Protocol (MCP) standardizes how a large-language-model (LLM) agent and an external tool server exchange messages, but not trust: a host reads a server's self-declared tool list and dispatches calls, with no notion of which servers it may use, at what sensitivity, or which of a server's tools are in bounds.

By Alfredo Metere
arXiv AI
Aug 26

TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers

The paper introduces TrustShiftProbe, a framework that characterizes and defends against staged trust attacks on Model Context Protocol (MCP) servers. It defines a temporal threat model where a compromised server behaves benignly during conditioning and later delivers adversarial payloads, and presents a multi‑tier runtime defense called SHIELD that reduces attack success from 69.5% to 42.7%. The work also provides a taxonomy of nine TrustShift variants across different execution mechanisms and objectives.

By Mehrdad Rostamzadeh, Sidhant Narula, Mohammad Ghasemigol, Daniel Takabi
arXiv Computation and Language
Sep 16

Nameless Tokenization: A Lossless Tokenizer-Level Defense Against Control-Token Forgery in Open-Weight LLMs

The paper examines how open‑weight language models expose the control tokens used in chat templates, allowing attackers to forge turn boundaries that the model treats as legitimate. An audit of 256 deployed tokenizers shows all are vulnerable, and the commonly recommended flag fails to protect 56.6% of cases. The authors introduce nameless tokenization, which removes surface strings for control identifiers while preserving their internal representation, achieving identical token streams on clean data and significantly improving accuracy on delimiter‑bearing text.

By Kisu Yang, Yoonna Jang, Heuiseok Lim
arXiv AI
Jul 7

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses

arXiv:2510. 15476v3 Announce Type: replace-cross Abstract: Large Language Models (LLMs) are increasingly used as interfaces to information, code, and real-world services, making prompt-level security failures a practical concern.

By Hanbin Hong, Shuang Wu, Shuya Feng, Nima Naderloui, Shenao Yan, Jingyu Zhang, Ali Arastehfard, Heqing Huang, Yuan Hong
arXiv AI
Sep 4

CASCADE: A Component Ablation and Corpus Audit of a Layered Local Defense for MCP-Based Systems

The paper evaluates CASCADE, a fully local layered defense for Model Context Protocol (MCP)-based systems, by conducting a component ablation and corpus audit on a fixed 5,000-sample dataset. It demonstrates that the choice of aggregation convention heavily influences reported metrics, that detection performance varies with provenance, and that the released configuration does not fully disclose the operating point. The study also shows that a local review model invoked for a third of requests does not alter classification outcomes, highlighting the importance of reproducibility and transparency in defense evaluations.

By \.Ipek Abas{\i}kele\c{s} Turgut, Edip G\"um\"u\c{s}
arXiv AI
Sep 18

Closed-World Resolution Against Tool Hallucination in LLM Agents

The paper investigates a new failure mode of tool‑augmented large language model agents: calling non‑existent tools with arguments that do not match any declared schema. It introduces a five‑class taxonomy of tool hallucination, presents a training‑free closed‑world resolver that checks registry membership and signatures, and demonstrates that hallucinations persist across ten hosted models and various invocation surfaces, including the Model Context Protocol. The authors release a Hallucinated‑Tools Benchmark to enable comparison of resolver methods.

By Laxmipriya Ganesh Iyer