arXiv AI

LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios

LPS-Bench is a benchmark designed to evaluate the safety awareness of computer‑use agents (CUAs) in long‑horizon planning tasks that involve tool workflows. It uses a template‑guided multi‑agent pipeline to generate user instructions, simulated toolkits, and case‑specific safety criteria, followed by human review, allowing scalable expansion without building separate application environments. The benchmark includes 570 cases from 65 scenarios across seven task domains and nine planning‑risk types, and an LLM‑based evaluator assesses tool choices, arguments, and responses throughout execution. Evaluations of 13 LLM agents show persistent safety failures in both benign and adversarial settings, with prompt‑based interventions providing only model‑dependent improvements.

arXiv AI
Aug 10

ForesightSafety-SAGE:A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents

arXiv:2606. 08531v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly evolving from simple text-based interaction systems into LLM agents that can maintain memory, use tools, access external environments, and execute tasks.

By Lu Jia, Haibo Tong, Feifei Zhao, Jindong Li, Dongqi Liang, Ping Wu, Qian Zhang, Yi Zeng
arXiv Machine Learning
Sep 22

Beyond Task Completion: Training Capable and Safe Computer-Use Agents

The paper introduces SCOPE, a method that post‑trains computer‑use agents to balance task completion with safety by conditioning actions on environmental risk. It combines supervised fine‑tuning on three trajectory types—capability demonstrations, safe continuations, and explicit refusals—followed by reinforcement learning to improve performance. Experiments starting from Qwen3.5‑9B show that SCOPE‑RL achieves high task success and attack‑avoidance rates, outperforming other agents on OSWorld and OS‑BLIND benchmarks.

By Zeyu Kang, Zhenyun Yin, Yang Zhang, Shan He, Shanzhe Lei, Yanjiu Zhong, Xinquan Chen, Yuhong Wang