arXiv AI

NEST: Nascent Encoded Steganographic Thoughts

arXiv:2602. 14095v2 Announce Type: replace Abstract: Monitoring chain-of-thought (CoT) reasoning is a foundational safety technique for large language model agents; however, this oversight is compromised if models learn to conceal their reasoning.

arXiv AI
3d ago

Learning Steganography Is Easy, Learning Steganographic Reasoning Is Hard

The paper investigates how large language models learn to hide their reasoning within text—termed steganographic reasoning—compared to related abilities like steganographic messaging and encoded reasoning. Across reinforcement learning, in-context learning, and supervised fine-tuning, models readily acquire messaging and encoded reasoning, but steganographic reasoning only emerges under supervised fine-tuning and requires substantially more training, unless a convenient cover task is provided. Even then, steganographic reasoning remains significantly harder than its neighboring capabilities.

By Julian Schulz, Lukas F\"ulle, Rieke Fruengel
arXiv Machine Learning
Sep 21

OverThink: Slowdown Attacks on Reasoning LLMs

The paper introduces OverThink, a slowdown attack that forces reasoning language models (RLMs) to produce many more reasoning tokens while still giving correct answers. By injecting decoy reasoning problems—such as Markov decision processes, language translation, or graphic comprehension—into the model’s context, attackers can dramatically increase token generation (up to 46× on SQuAD and 17× on coding agents). The study evaluates the attack on both proprietary and open-source RLMs across multiple datasets, explores multimodal and coding‑agent variants, and tests several defenses, concluding that defending against OverThink is challenging and that newer RLMs are even more vulnerable due to higher per‑token costs and increased reasoning token usage.

By Abhinav Kumar, Jaechul Roh, Ali Naseh, Marzena Karpinska, Mohit Iyyer, Amir Houmansadr, Eugene Bagdasarian
arXiv AI
Aug 24

Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification

The paper identifies a vulnerability in large language models where harmful intent can be hidden within benign narratives, a phenomenon termed Semantic Camouflage. By examining latent activation patterns across several small language model families, the authors discover an "Intent Horizon"—a layer depth where harmful intent representations collapse. They propose Latent Intent Verification (LIV), a lightweight probing defense that detects harmful intent in early layers and outperforms existing guardrails on the PKU-SafeRLHF dataset.

By Md. Hasib Ur Rahman
arXiv AI
Sep 18

FORGE: Forensic Reasoning with Grounded Evidence

FORGE is a forensic deepfake analysis system that provides region‑grounded natural language explanations for image manipulations. It addresses the inductive bias mismatch of multimodal large language models by adding a Vision‑Only Model trained on dense patch prediction, allowing the language model to interleave tokens with preserved spatial correspondence. Across face‑manipulated and fully synthetic content, FORGE delivers fine‑grained attribute queries and outperforms in‑domain baselines, with region‑specific evaluation and human studies confirming explanation faithfulness.

By Rohit Kundu, Shan Jia, Vishal Mohanty, Athula Balachandran, Amit K. Roy-Chowdhury