arXiv:2609.21841v1 Announce Type: new
Abstract: Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically val...
By Abbas Raza Ali, Muhammad Ajmal Siddiqui, Moona Zahid
arXiv:2606. 15563v1 Announce Type: new Abstract: AI systems increasingly delegate decisions to specialized models, evaluators, tools, and supervisory controllers.
By Carlos R. B. Azevedo
arXiv:2607. 13034v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task actually requires.
By Junjie Yin, Xinyu Feng
arXiv:2609.37501v1 Announce Type: cross
Abstract: We propose RegLLM, a diagnostic harness for bounded autonomy in regulated agentic workflows. It instruments six trustworthiness signals: citation val...
By Dipankar Sarkar
arXiv:2605. 25739v2 Announce Type: replace Abstract: We prove that no reinforcement learning policy with confidence-gated autonomy can simultaneously achieve maximum helpfulness, optimal calibration, and full autonomy under rational oversight, whenever some tasks exceed the agent's reliable competence: the Behavioral Credibility Trilemma.
By Lauri Lov\'en, Nam Do, Hassan Mehmood, Dinesh Kumar Sah, Sasu Tarkoma
RGDT-Bench is a new benchmark that evaluates large language models on Rule‑Governed Decision Tasks, where models must apply external rules to facts, justify decisions, and provide checkable justifications. The benchmark offers 202.1K condition‑level supervision slots across four task tracks and eight task‑probe combinations, and it labels warrant completeness through label‑blind extraction and deterministic checks. Evaluation shows that among correct responses, 40.2% of warrants are incomplete, and existing evaluators struggle to detect this, prompting the authors to train a reward model that improves AUROC to 69.24% and outperforms outcome‑supervised baselines.
arXiv:2608. 16402v1 Announce Type: new Abstract: Large language model-based agentic frameworks primarily optimize capability: whether an agent can reason, retrieve information, call tools, delegate work, and complete a goal.
By Bhaskar Tripathi, Anurag Kumar, Ramendra Kumar, Bhavesh Gadhe
The paper introduces the concept of LLM Parkinsonism, describing how large language models can persist in low‑value actions after completing their objectives. It proposes a Global Executive Control (GEC) architecture that separates action generation from project‑level oversight, achieving comparable success to candidate‑set control while significantly reducing token usage and complexity. Experimental results on a 24,000‑episode benchmark show GEC cuts mean token use by 36.4% and limits token consumption at the 40,000‑token ceiling by 18.7%, eliminating pre‑completion drift.
By Dongsheng Xiao, Zeyuan Wang, Xuzhe Xia, Bo Zhao, Yankai Cao
The paper proposes a metric framework to differentiate cognitive amplification—where AI enhances human performance without eroding human capability—from cognitive delegation, which relies heavily on AI reasoning. It introduces four metrics (CAI*, D, HRI, HCDR) and tests them in NetLogo simulations across various reliance and dependency scenarios. The results show that positive collaborative gain is only achievable when an explicit interaction term is added, indicating that mere prevention of capability erosion is insufficient for genuine amplification.
By Eduardo Di Santi, Carla Florida
arXiv:2608. 11727v1 Announce Type: new Abstract: When a coding agent obeys a rule, it may simply have been going to do that anyway.
By Zining Huang, Haoran Que, Hong Zeng, Ge Zhang, Zuo Wang, Jin Chen, Haodong Wang, Zhongfei Hou, Changxin Pu, Shen Yan, Wenhao Huang
arXiv:2602. 13213v2 Announce Type: replace Abstract: Commercial insurance underwriting is a labor-intensive process that requires manual review of extensive documentation to assess risk and determine policy pricing.
By Joyjit Roy, Samaresh Kumar Singh
arXiv:2607. 08964v1 Announce Type: new Abstract: AI agents have become capable of autonomously completing short, well-specified tasks.
By Zongxia Li, Zhongzhi Li, Yucheng Shi, Ruhan Wang, Junyao Yang, Zhichao Liu, Xiyang Wu, Anhao Li, Yue Yu, Ninghao Liu, Lichao Sun, Haotao Mi, LeoweiLiang