arXiv AI By An Luo, Jie Ding

Nonuniformity Principle in Human-AI Coworking

Read the original on arXiv AI →

arXiv:2607. 16530v1 Announce Type: new Abstract: As generative AI is increasingly applied to automate multi-step and high-stake workflows, human judgment and involvement remain essential for ensuring the quality of AI-generated outputs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

AI Agents Push Humans Out of the Loop

AI agents are increasingly autonomous, posing significant risks that current designs hinder effective human oversight. The paper argues that oversight is degraded by both design choices and the cognitive decline of users who rely heavily on automation. It calls for prioritizing human cognitive needs in AI agent development, proposing design affordances and protocols to maintain critical judgment and counter skill atrophy.

By Margaret Mitchell, Avijit Ghosh, Samir Passi
arXiv Computer Vision
Aug 28

OS-Marathon: Benchmarking Computer-Use Agents on Vast-Horizon, Repetitive Tasks

OS-Marathon is a new benchmark that tests computer‑use agents on vast‑horizon, repetitive tasks, covering 100 tasks across five scenarios and ten domains. The study shows that current state‑of‑the‑art agents perform poorly on these tasks, and that simply decomposing workflows into subtasks does not solve the problem. Introducing a cost‑friendly personalization method called GraphDemo, which adapts agents from a single human demonstration, improves performance, highlighting the value of human guidance for these challenging tasks.

By Jing Wu, Wenjie Ai, Daphne Barretto, Yiye Chen, Qingyu Chen, Yuhang He, Pranit Chawla, Nicholas Gyd\'e, Yanan Jian, Vibhav Vineet