This survey reviews 148 linguistic steganographic methods, 60 countermeasures, 23 evaluation metrics, and 9 open challenges, providing taxonomies, reviews, and adoption analyses. It identifies five paradigm shifts brought by large language models: moving from covertext modification to prompt-only generation, from heuristic to provable security, from white-box symmetric models to black-box or asymmetric access, from security-centric designs to joint optimization, and from text-quality concerns to engineering issues. The paper aims to serve as a reference and roadmap for practical and responsible linguistic steganography in the LLM era.
By Ruiyi Yan, Chenhui Chu, Zhongliang Yang, Yugo Murawaki
arXiv:2606. 09411v1 Announce Type: cross Abstract: Large language models can be fine-tuned to encode prompt-borne secrets into fluent, seemingly benign outputs.
By Charles Westphal, Timothy Douglas, Keivan Navaie, Tiago Pimentel, Fernando E. Rosas
arXiv:2601. 22818v2 Announce Type: replace-cross Abstract: Fine-tuned LLMs can covertly encode prompt secrets into outputs via steganographic channels.
By Charles Westphal, Keivan Navaie, Fernando E. Rosas
WeaveMark is a new multi‑bit watermarking scheme for large language models that improves payload capacity, extraction accuracy, and text quality by using coded payload spreading, soft‑decision error‑correcting codes, and unbiased multilayer reweighting. It also adds zero‑bit layers for reliable detection of watermark presence. Experiments demonstrate significant gains, achieving an 89.8% match rate for 32‑bit messages at 200 tokens and maintaining 86.0% accuracy under 10% substitution attacks on 16‑bit messages, far outperforming the BiMark baseline.
By Gang-Hyun Park, Ju-Hyeong Lee, Hee-Youl Kwak, Dae-Young Yun
arXiv:2607. 05353v1 Announce Type: cross Abstract: Watermarking methods embed imperceptible and verifiable signals into text generated by large language models (LLMs).
By Xuyang Chen, Xiang Li, Yangxinyu Xie, Qi Long
arXiv:2512. 13325v2 Announce Type: replace-cross Abstract: Securing digital text is becoming increasingly relevant due to the widespread use of large language models.
By Malte Hellmeier
arXiv:2606. 09135v1 Announce Type: cross Abstract: We demonstrate that widely deployed Large Language Model (LLM) inference stacks harbor a steganographic channel that requires no modification to model weights, sampling code, or output distributions.
By Felix M\"achtle, Jonas Sander, Sebastian Berndt, Ben Weimar, Nils Loose, Thomas Eisenbarth
arXiv:2609.22392v1 Announce Type: new
Abstract: Image steganography hides secret message within normal images, with most existing works relying on cover-preserving transmission. However, such a parad...
By Qi Li, Jidong Yang, Huaike Yu, Chunpeng Wang, Suo Gao, Herbert Ho-Ching Iu, Yuantian Miao, Bin Ma, Xiao Chen
The paper introduces CARTS, a steganographic method that uses autoregressive language models to encode a payload text into a stegotext of identical token length by preserving per‑position rank information across contexts. It provides a formal security analysis, proving exact correctness under deterministic model assumptions, and defines key security notions such as context search, key collisions, message equivocation, and non‑commutativity of encoding maps. Empirical tests on Llama 3 8B confirm perfect payload recovery, no random key collisions, and no commuting key pairs, indicating resistance to the studied attack vectors.
By Wissam Ghantous, Alexander V. Mantzaris
arXiv:2606. 28425v1 Announce Type: cross Abstract: Increasingly autonomous agentic AI systems pose novel multi-agent risks, such as secret collusion via covert communication channels.
By Jimmy Laurence Rippin, Simon C. Marshall, David Demitri Africa, Christian Schroeder de Witt
arXiv:2610.01871v1 Announce Type: cross
Abstract: Multimodal Retrieval-Augmented Generation (MRAG) has emerged as a reliable and cost-effective technique of grounding the generative capabilities of M...
By Maria Carmen Jica, Ali Satvaty, Suzan Verberne, Fatih Turkmen
arXiv:2608.27899v1 Announce Type: cross
Abstract: With the growing prevalence of large language model (LLM) generated content, watermarking is considered a promising approach for attributing text to...
By Miroojin Bakshi, Saksham Rastogi, Danish Pruthi