arXiv AI By Wenhao Lin, Chenyu Yu, Xingwei Lin, Sicong Cao, Xiang Chen, Lei Xue, Le Yu, Letian Sha, Chunming Wu

DreamGuard: Efficient Runtime Guardrail for LLM Agents via Risk-Aware World Model

Read the original on arXiv AI →

arXiv:2608. 05695v1 Announce Type: new Abstract: As large language model (LLM) agents increasingly invoke external tools and interact with real-world systems, unsafe actions may cause irreversible consequences on external states, user data, and downstream services.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 6

From Risk Classification to Action Plan Remediation: A Guardrail Feedback Driven Framework for LLM Agents

arXiv:2606. 05805v1 Announce Type: new Abstract: LLM-based guardrails typically safeguard agents by evaluating proposed actions or inputs before execution, producing safety signals such as binary allow/deny decisions, risk categories, and/or explanatory rationales about potential policy violations.

By Yuhao Sun, Jiacheng Zhang, Shaanan Cohney, Zhexin Zhang, Feng Liu, Xingliang Yuan
arXiv AI
Sep 17

HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark

HINTBench is a new benchmark for evaluating agents’ intrinsic risk, comprising 596 trajectories (400 synthetic risky, 136 synthetic safe, 30 real risky, 30 real safe) with an average length of 24 steps. It supports three tasks—risk detection, risk-step localization, and intrinsic failure-type identification—using a unified five-constraint taxonomy. Experiments show a large performance gap: while large language models can detect risky trajectories, they score below 37 on strict-F1 for risk-step localization, and existing guard models transfer poorly to this setting.

By Jiacheng Wang, Jinchang Hou, Fabian Wang, Ping Jian, Chenfu Bao, Zhonghou Lv