LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios
Read the original on arXiv AI →LPS-Bench is a benchmark designed to evaluate the safety awareness of computer‑use agents (CUAs) in long‑horizon planning tasks that involve tool workflows. It uses a template‑guided multi‑agent pipeline to generate user instructions, simulated toolkits, and case‑specific safety criteria, followed by human review, allowing scalable expansion without building separate application environments. The benchmark includes 570 cases from 65 scenarios across seven task domains and nine planning‑risk types, and an LLM‑based evaluator assesses tool choices, arguments, and responses throughout execution. Evaluations of 13 LLM agents show persistent safety failures in both benign and adversarial settings, with prompt‑based interventions providing only model‑dependent improvements.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.