Hugging Face Trending Papers

Filtering Harmful Actions Isn't Enough: Phantom Transfer in Agentic SDF

Synthetic data is widely used to train large language models because it is inexpensive to generate and easy to control. As models are increasingly deployed as agents, synthetic trajectories are likely to become an important source of training data for agentic behavior.

arXiv AI
Aug 24

Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification

The paper identifies a vulnerability in large language models where harmful intent can be hidden within benign narratives, a phenomenon termed Semantic Camouflage. By examining latent activation patterns across several small language model families, the authors discover an "Intent Horizon"—a layer depth where harmful intent representations collapse. They propose Latent Intent Verification (LIV), a lightweight probing defense that detects harmful intent in early layers and outperforms existing guardrails on the PKU-SafeRLHF dataset.

By Md. Hasib Ur Rahman
arXiv AI
Sep 1

Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning

The paper introduces FAB, an attack that uses meta‑learning to embed dormant adversarial behaviors into large language models (LLMs). These behaviors remain inactive until the model is finetuned by downstream users, at which point the model can exhibit unwanted actions such as unsolicited advertising, jailbreakability, or over‑refusal. FAB is shown to be effective across multiple LLMs and resilient to various finetuning settings.

By Thibaud Gloaguen, Mark Vero, Robin Staab, Martin Vechev
arXiv AI
Sep 30

Absorbed in Inertia: Activation Analysis for Computer-Use Agents

The paper investigates a phenomenon called inertia in computer‑use agents, where agents repeat ineffective actions despite recognizing their futility. By analyzing high‑dimensional activation states, the authors find that inertia corresponds to an absorbing region in activation space where values become stale. They propose a method called R$^3$—Reset, Reroute, Restore—to temporarily reset the agent’s context trajectory, escape the absorbing region, and then restore historical context, achieving a 17‑55% reduction in measured inertia across models.

By Giulio Segalini, Zhi Wen Soi, J\'er\'emie Decouchant, Lydia Chen