arXiv AI By Zhonghao Zhan, Xiao Ma, Hamed Haddadi

Safe Skill Retirement for Physical Agents

Read the original on arXiv AI →

The paper introduces a method for safely retiring procedural guidance in AI agents that control physical actions. It proposes matched authority counterfactuals and a two‑gate retirement certificate to ensure that reductions preserve authorized utility while eliminating unauthorized protected effects. Experiments across multiple models and skill bundles show that task‑certified reductions can remove most skill clauses, but only a combined protocol passes both utility and safety gates in all tested configurations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 25

Who Holds the Pen? Let Specifications, Not Agents, Sign Off

The paper argues that large language model agents should not be the sole authority on whether they have satisfied a task. It identifies two gaps—understanding–execution and state–authority—where agents may claim completion without actually meeting the specification. The authors propose SpecHarness, a framework that separates agent proposals from authoritative state by requiring evidence from qualified providers to confirm compliance, and demonstrate its effectiveness on guideline‑following and artifact‑generation tasks.

By Haiqing Li, Xin Ma, Yinhao Wu, Wenliang Zhong, Feng Jiang, Thao M. Dang, Xiao Hu, Hehuan Ma, Yuzhi Guo, Junzhou Huang
arXiv AI
Aug 20

Task-Conditioned Least-Privilege Learning for Executable Terminal and MCP Agents

The paper introduces a post‑training framework that teaches a 4B‑parameter language model to exercise task‑conditioned authority in executable terminal and Model Context Protocol (MCP) environments. By auditing each action across six risk dimensions with deterministic verifiers and optimizing for task‑specific excess‑privilege values, the authors achieve 98.48% safe success and reduce excess‑authority errors from 4.56% to 0.79% on held‑out tasks. The study also demonstrates capability retention, prompt‑directed improvement, and generalization over a 400‑task continuation test.

By Alexander Tu, Michael Tu