AURA-Eval: Evaluation Framework for Acting Under Risk Awareness in LLM Agent Trajectories
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2604. 02022v4 Announce Type: replace Abstract: Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step interactions rather than isolated prompts or final responses.
arXiv:2606. 09890v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as autonomous agents capable of executing multi-step action trajectories toward a given objective.
arXiv:2606. 08531v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, access external environments, and execute tasks.
arXiv:2607. 29254v1 Announce Type: new Abstract: AI agents extend large language models (LLMs) with external tools, enabling them to perform complex tasks and translate model outputs into consequential real-world actions.
HINTBench is a new benchmark for evaluating agents’ intrinsic risk, comprising 596 trajectories (400 synthetic risky, 136 synthetic safe, 30 real risky, 30 real safe) with an average length of 24 steps. It supports three tasks—risk detection, risk-step localization, and intrinsic failure-type identification—using a unified five-constraint taxonomy. Experiments show a large performance gap: while large language models can detect risky trajectories, they score below 37 on strict-F1 for risk-step localization, and existing guard models transfer poorly to this setting.
arXiv:2606. 08531v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, access external environments, and execute tasks.