AI agents are increasingly autonomous, posing significant risks that current designs hinder effective human oversight. The paper argues that oversight is degraded by both design choices and the cognitive decline of users who rely heavily on automation. It calls for prioritizing human cognitive needs in AI agent development, proposing design affordances and protocols to maintain critical judgment and counter skill atrophy.
By Margaret Mitchell, Avijit Ghosh, Samir Passi
arXiv:2606. 05770v1 Announce Type: cross Abstract: AI is changing how software engineers work, but it often comes with hidden burdens and costs.
By Vahid Garousi
arXiv:2608. 14870v1 Announce Type: new Abstract: Long-horizon agents can execute continuously, but human attention remains intermittent and scarce.
By Chen Chen, Zhehuai Chen
OS-Marathon is a new benchmark that tests computer‑use agents on vast‑horizon, repetitive tasks, covering 100 tasks across five scenarios and ten domains. The study shows that current state‑of‑the‑art agents perform poorly on these tasks, and that simply decomposing workflows into subtasks does not solve the problem. Introducing a cost‑friendly personalization method called GraphDemo, which adapts agents from a single human demonstration, improves performance, highlighting the value of human guidance for these challenging tasks.
By Jing Wu, Wenjie Ai, Daphne Barretto, Yiye Chen, Qingyu Chen, Yuhang He, Pranit Chawla, Nicholas Gyd\'e, Yanan Jian, Vibhav Vineet
Scaling human oversight of AI systems for tasks that are difficult to evaluate.
arXiv:2510. 26518v2 Announce Type: replace Abstract: Human feedback is critical for aligning AI systems to human values.
By Rishub Jain, Sophie Bridgers, Lili Janzer, Rory Greig, Tian Huey Teh, Vladimir Mikulik