arXiv AI

Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw

The paper examines how users delegate tasks to the AI agent OpenClaw by analyzing 73,093 Reddit posts. It identifies 21 human values grouped into six categories—such as Autonomous Operation, Dependable Operation, Affordable Operation, Bounded Reach, Reviewability, and Equitable Access—and finds that values are largely satisfied when users describe the agent’s outputs but often unmet when users discuss supervising the agent. The authors term this pattern "value‑sensitive delegation," emphasizing that supporting human values requires attention to both what an agent does and the conditions users set around its use.

arXiv AI
Aug 20

Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025

The paper titled "Professional Software Developers Don't Vibe, They Control: AI Agent Use for Coding in 2025" examines how experienced developers employ AI agents in software development. Through field observations and surveys, it finds that developers value agents for productivity but maintain control over design and implementation to ensure quality. They use agents as collaborative tools rather than full delegation, selecting tasks based on suitability and leveraging their expertise to guide agent behavior.

By Ruanqianqian Huang, Avery Reyna, Sorin Lerner, Haijun Xia, Brian Hempel
arXiv AI
Sep 4

Value-Preserving Architectures for Agentic AI Systems

The paper "Value-Preserving Architectures for Agentic AI Systems" discusses how architectural choices in large language model-based multi‑agent systems (MAS) can promote human‑centered values such as privacy, fairness, and safety. It introduces three value‑preserving architectural patterns: a privacy‑aware federated topology, a distributed architecture that encourages pluralism and diversity, and a guard‑agent design to detect and mitigate unfairness. Representative use cases illustrate how these patterns can be applied in real‑world scenarios, aiming to provide guidelines for building trustworthy MAS.

By Alessandro Pesare, Tommaso Dolci, Katja Hose, Emanuel Sallinger
arXiv AI
Jun 12

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility

arXiv:2606. 13608v1 Announce Type: new Abstract: Agent systems are advancing quickly across domains, but their evaluation remains fragmented.

By Xiaoyuan Liu, Jianhong Tu, Yuqi Chen, Siyuan Xie, Sihan Ren, Tianneng Shi, Gal Gantar, Evan Sandoval, Donghyun Lee, Daniel Miao, Peter J. Gilbert, Nick Hynes, Mauro Staver, Warren He, David Marn, Andrew Low, Xi Zhang, Elron Bandel, Michal Shmueli-Scheuer, Siva Reddy, Alexandre Drouin, Alexandre Lacoste, Ramayya Krishnan, Elham Tabassi, Yu Su, Victor Barres, Chenguang Wang, Wenbo Guo, Dawn Song
arXiv AI
Sep 12

Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks

The article "Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks" surveys the lack of a standard definition for AI agents and organizes this ambiguity into five dimensions: environmental interaction, learning and adaptation, autonomy, goal‑directed behavior, and temporal coherence. It reviews how each dimension has been conceptualized in prior work and compiles the metrics, benchmarks, and evaluation frameworks used to assess them. The authors also introduce the Agent Compendium, a public digital resource that extends these evaluation methods, aiming to provide a common structure for evaluating and comparing agent capabilities across AI systems.

By Mia Lassiter, Brinnae Bent