arXiv AI By Charles Westphal, Keivan Navaie, Fernando E. Rosas

Hide and Seek in Embedding Space: Geometry-based Steganography and Detection in Large Language Models

Read the original on arXiv AI →

arXiv:2601. 22818v2 Announce Type: replace-cross Abstract: Fine-tuned LLMs can covertly encode prompt secrets into outputs via steganographic channels.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 30

ToxScreen: Detecting Whether an LLM Has Been Poisoned

arXiv:2607. 26849v1 Announce Type: cross Abstract: As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time.

By Anthony Hughes, Nicole Xing, Collin Francel, Andy Kim, Andrew Draganov