arXiv AI By Shiva Pochampally, Shengwei An, Yan Chen

Assistant or Actor? Student Trust, Control, and Delegation Regret When Using a General-Purpose AI Agent

Read the original on arXiv AI →

arXiv:2607. 18257v1 Announce Type: cross Abstract: When AI agents shift from answering questions to taking actions, users face a new problem: deciding what to delegate, to a system whose action space they cannot fully anticipate.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 25

When Agents Act Unwatched: The Reduced-Supervision Paradox in Agentic AI

The paper "When Agents Act Unwatched: The Reduced‑Supervision Paradox in Agentic AI" discusses how the promise that AI systems will continue acting after users stop watching creates an accountability inversion. It argues that as stepwise supervision recedes, verification shifts into the runtime infrastructure—authority, records, interrupts, outcome checks, and repair—forming what the authors call the reduced‑supervision paradox. A 63‑artifact audit across research papers and engineering sources shows that agents’ action surfaces are more visible than the mechanisms needed to hold them accountable, with tool mediation and monitoring traces appearing in 40 and 37 artifacts, while checkpoint placement, validator independence, recovery, and contestability are rarely visible. "whyItMatters":"The study highlights that observable action paths can replace accountability when verification is moved onto users after meaningful intervention is no longer possible."

By Hanjing Shi, Dominic DiFranzo
arXiv AI
Aug 20

Task-Conditioned Least-Privilege Learning for Executable Terminal and MCP Agents

The paper introduces a post‑training framework that teaches a 4B‑parameter language model to exercise task‑conditioned authority in executable terminal and Model Context Protocol (MCP) environments. By auditing each action across six risk dimensions with deterministic verifiers and optimizing for task‑specific excess‑privilege values, the authors achieve 98.48% safe success and reduce excess‑authority errors from 4.56% to 0.79% on held‑out tasks. The study also demonstrates capability retention, prompt‑directed improvement, and generalization over a 400‑task continuation test.

By Alexander Tu, Michael Tu
arXiv AI
2d ago

Cybernetic and Epistemic: A Missing Vocabulary for Trustworthy Agentic Delegation

The paper argues that as AI systems increasingly generate code, the bottleneck has shifted to supervising these systems, revealing a vocabulary gap between cybernetic coordination (actions aligning with the world) and epistemic coordination (understanding that can be verified). It critiques current oversight that merely approves outputs, proposing instead that every consequential choice by an agent must include a retrievable condition explaining why it was made, enabling third‑party verification. The authors illustrate this with three delegation episodes, introduce a two‑part reconstruction test, and propose the ORRCF convention to embed such conditions in all recorded decisions.

By J\'er\'emie Lumbroso