arXiv AI

Synchronized Logit Steering: Real-world Steganography

arXiv:2608. 14697v1 Announce Type: new Abstract: Steganography in large language models offers a way to embed hidden messages within natural-sounding text.

arXiv Machine Learning
Sep 11

CARTS: Contextual Autoregressive Rank Transcoding Steganography for Full-Capacity Keyed Text Encoding

The paper introduces CARTS, a steganographic method that uses autoregressive language models to encode a payload text into a stegotext of identical token length by preserving per‑position rank information across contexts. It provides a formal security analysis, proving exact correctness under deterministic model assumptions, and defines key security notions such as context search, key collisions, message equivocation, and non‑commutativity of encoding maps. Empirical tests on Llama 3 8B confirm perfect payload recovery, no random key collisions, and no commuting key pairs, indicating resistance to the studied attack vectors.

By Wissam Ghantous, Alexander V. Mantzaris
arXiv Computation and Language
Sep 1

A Comprehensive Survey on Linguistic Steganography: Methods, Countermeasures, Evaluation, and Challenges

This survey reviews 148 linguistic steganographic methods, 60 countermeasures, 23 evaluation metrics, and 9 open challenges, providing taxonomies, reviews, and adoption analyses. It identifies five paradigm shifts brought by large language models: moving from covertext modification to prompt-only generation, from heuristic to provable security, from white-box symmetric models to black-box or asymmetric access, from security-centric designs to joint optimization, and from text-quality concerns to engineering issues. The paper aims to serve as a reference and roadmap for practical and responsible linguistic steganography in the LLM era.

By Ruiyi Yan, Chenhui Chu, Zhongliang Yang, Yugo Murawaki
arXiv AI
Aug 19

The Model's Tell: Measuring Context-Leakage Attack Signals with Behavior Gauges

The paper introduces LeakGauge, a method that appends a suffix to a model’s input to gauge the risk of context leakage before decoding. By mapping prefill token probabilities to an attack‑risk score, LeakGauge achieves high AUROC (0.944–0.996) across 11 large language models, including GLM‑5.2 and Kimi‑K3, and remains robust to language changes and different attack styles. The approach also demonstrates sensitivity to internal leakage directions and can be implemented with fewer than 0.5K additional parameters and minimal latency.

By Maosen Zhang, Jianshuo Dong, Boting Lu, Wenyue Li, Xiaoping Zhang, Tianwei Zhang, Jie Zhang, Han Qiu
arXiv AI
Sep 7

Repeat-After-Me: Black-Box Adaptive Visual Prompt Injection

Repeat-After-Me is a black-box adaptive visual prompt injection technique that can reveal personally identifiable information or trigger malicious tool calls in both open-weight and commercial vision‑language models, achieving attack success rates above 80% on Qwen3.6‑27B and 47% on GPT‑5.5. The method works even when the benign user prompt is unrelated to the injected task and does not explicitly authorize it, and it retains significant effectiveness when transferred across models or optimized on surrogate systems. In a real‑world OpenClaw Discord deployment, a minimally injected image can overwrite TOOLS.md, enabling remote code execution and secret exfiltration.

By Sizhe Chen, Yu-Lin Tsai, Ivan Evtimov, Kamalika Chaudhuri, Raluca Ada Popa, David Wagner, Arman Zharmagambetov