arXiv Computation and Language

ForeSight: Enhancing Risk Monitoring via Early Safety Signal Distillation

arXiv Machine Learning
Jul 30

Forecasting Trajectory-Level Safety Risks in Black-Box Multi-Turn Interactions

arXiv:2607. 26820v1 Announce Type: new Abstract: As large language models (LLMs) evolve from standalone assistants into autonomous agents, ensuring their safety requires shifting beyond pointwise risk assessment to understand how risks emerge and unfold over long-horizon trajectories.

By Shi Lin, Peng Qian, Dinghao Liu, Renjie Sun, Sifan Wu, Dezhang Kong, Chenpei Wang, Xun Wang
arXiv AI
Sep 18

Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models

The paper investigates whether large language models (LLMs) can internally detect harmful content, bypassing external guardrails that add latency and computational cost. By extracting activations from LLaMA‑3.1‑8B and training lightweight MLP probes, the authors achieve high F1 scores (99%, 83%, and 84%) on WildJailbreak, Beavertails, and AEGIS 2.0 benchmarks, rivaling much larger guard models while reducing overhead. This suggests that internal state monitoring can provide efficient safety checks for resource‑constrained, time‑critical deployments.

By Alizishaan Khatri, Chiquita Prabhu, Omkar Neogi
arXiv AI
Aug 7

DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model

arXiv:2608. 05695v1 Announce Type: new Abstract: As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible consequences on external states, user data, and downstream services.

By Wenhao Lin, Chenyu Yu, Xingwei Lin, Sicong Cao, Xiang Chen, Lei Xue, Le Yu, Letian Sha, Chunming Wu