arXiv:2601. 22818v2 Announce Type: replace-cross Abstract: Fine-tuned LLMs can covertly encode prompt secrets into outputs via steganographic channels.
By Charles Westphal, Keivan Navaie, Fernando E. Rosas
This survey reviews 148 linguistic steganographic methods, 60 countermeasures, 23 evaluation metrics, and 9 open challenges, providing taxonomies, reviews, and adoption analyses. It identifies five paradigm shifts brought by large language models: moving from covertext modification to prompt-only generation, from heuristic to provable security, from white-box symmetric models to black-box or asymmetric access, from security-centric designs to joint optimization, and from text-quality concerns to engineering issues. The paper aims to serve as a reference and roadmap for practical and responsible linguistic steganography in the LLM era.
By Ruiyi Yan, Chenhui Chu, Zhongliang Yang, Yugo Murawaki
arXiv:2606. 28425v1 Announce Type: cross Abstract: Increasingly autonomous agentic AI systems pose novel multi-agent risks, such as secret collusion via covert communication channels.
By Jimmy Laurence Rippin, Simon C. Marshall, David Demitri Africa, Christian Schroeder de Witt
arXiv:2602. 14095v2 Announce Type: replace Abstract: Monitoring chain-of-thought (CoT) reasoning is a foundational safety technique for large language model agents; however, this oversight is compromised if models learn to conceal their reasoning.
By Artem Karpov
arXiv:2606. 09135v1 Announce Type: cross Abstract: We demonstrate that widely deployed Large Language Model (LLM) inference stacks harbor a steganographic channel that requires no modification to model weights, sampling code, or output distributions.
By Felix M\"achtle, Jonas Sander, Sebastian Berndt, Ben Weimar, Nils Loose, Thomas Eisenbarth
arXiv:2511.14301v4 Announce Type: replace-cross
Abstract: Transformer-based models are highly susceptible to backdoor attacks via supervised fine-tuning (SFT). To red-team existing data-poisoning def...
By Eric Xue, Ruiyi Zhang, Pengtao Xie