HARDE: Optimizing Agent Harnesses for Runtime Risk Detection and Execution Control
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
HarnessRisk is a lifecycle-oriented benchmark for evaluating safety in agent harnesses that manage tools, extensions, state, permissions, and external actions. It defines six operational phases—Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and Incident Recovery—and includes 128 sandboxed cases pairing benign user objectives with adversarial instructions. Across three harnesses, six language models, and 14 configurations, attack success rates vary from 12.6% to 80.9%, with the most vulnerable phase being Harness Configuration. "whyItMatters":"The benchmark demonstrates that safety failures can arise in multiple harness responsibilities and that even explicit risk detection does not guarantee safe action, underscoring the need for comprehensive evaluation across model and harness configurations."
arXiv:2608. 12977v1 Announce Type: cross Abstract: The expanding operational capabilities of large language model (LLM) agents introduce sophisticated security threats.
SafeCoEvo is a test‑time framework that co‑evolves safety harnesses and guards for large language model agents. It uses a short‑term S‑Harness to quickly externalize recent runtime experience into explicit safety knowledge, and a long‑term GuardVPO to internalize accumulated experience into parametric risk‑judgment capabilities. This dual adaptation improves safety and task success, reducing unsafe outcomes by 10.05% and increasing task success by 12.15% over the strongest baseline.
arXiv:2609.14987v1 Announce Type: cross Abstract: Large language model (LLM) agents interact with external environments through tool invocation, but tool outputs can also expose them to indirect prom...
arXiv:2609.15134v1 Announce Type: new Abstract: Computer-use agents increasingly interact with browsers, terminals, file systems, and external services, introducing safety risks that emerge through r...
arXiv:2607. 01793v1 Announce Type: new Abstract: LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks.