arXiv AI

Strategic Decision Support for AI Agents

arXiv:2606. 12587v1 Announce Type: new Abstract: Traditionally, decision support studies how humans use machine learning models to make better decisions.

arXiv AI
Sep 25

Human-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation

The paper introduces a framework for hypothesis testing that combines inexpensive AI judgments with selective human verification to control type‑I and type‑II errors while minimizing cost. It derives an information‑theoretic lower bound on the minimum cost and proposes the SCALE policy, a sequential, cost‑aware strategy that adapts AI scoring and human escalation. SCALE is proven valid for finite samples and asymptotically matches the lower bound, achieving significant savings when both AI and human inputs are valuable.

By Dae Woong (David), Ham, Xuejun Zhao, Stefanus Jasin, Fenghua Yang
arXiv AI
Sep 10

Beyond "AI Helps Humans": Decision-Targeted Evaluation Design for Human-Agent Teams in the Agentic Era

The paper introduces TEAM-Design, a rule that assigns two replay probabilities to each task—one for a human-only replay and one for an agent-only replay—based on how difficult it is to predict the missing baseline outcome and the cost of replay. It addresses the challenge of deciding whether to keep a human-AI workflow or replace it with a single actor when only one outcome can be observed after deployment. The authors prove that TEAM-Design solves the budgeted design problem and controls error rates, and demonstrate its effectiveness on clinical and coding benchmarks, noting it excels when one comparison is clearly harder than the other.

By Hamed Khosravi, Xiaoming Huo
arXiv AI
Aug 28

Don't Overthink, Don't Underthink: Toward Adaptive Reasoning in Agentic AI

The paper argues that large language models need adaptive reasoning rather than fixed reasoning budgets. It shows that over‑reasoning leads to high computational cost without accuracy gains, while under‑reasoning results in incorrect or incomplete solutions. The authors evaluate these failure modes on MATH‑500 and the GAIA benchmark, highlighting the need for dynamic reasoning allocation in agentic AI systems.

By Md Jueal Mia, M. Hadi Amini
arXiv AI
Sep 3

Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality

The paper introduces ProSE, a framework for AI assistants that generate proposals while considering users’ bounded rationality and evaluability constraints. It proposes a KL‑regularised bounded‑rational binary response model and a depth‑2 Bayes‑adaptive planner, “ProSE‑Plan,” which scores proposals by expected responses and resulting belief updates. Experiments on graph simulations show that “ProSE‑Plan” outperforms evaluability‑unaware and myopic baselines, especially when evaluation cost is high, and that informative probes are crucial for effective assistance.

By Yifan Zhu, Sammie Katt, Samuel Kaski
arXiv AI
Sep 10

Agentic ML Exploration (A-MLE) for Ads Ranking

arXiv:2609.08248v1 Announce Type: new Abstract: Modern industrial ads ranking stacks are increasingly bottlenecked not by model capacity or training compute, but by the throughput of human ML iterati...

By Erwin Gao, Vinodh Kumar Sunkara, Jingyi Guan, Qinjin Jia, Hangjun Xu, Xiang Ji, Sherman Wong, Surya Teja Chavali, Pratik Vaishnavi, Aryan Pandhi, Xiaoyu Deng, Zhaodong Wang, Samarth Inani, Fan Yang, Jakob Moberg, Zoe Zu, Nicolas Bievre, Sami Khenissi, Amit Jaspal, Ehsan Fakharizadi, Srinidhi Viswanathan, Dorothy Sun, Abishek Vanam, Sneha Iyer, Sheela Yadawad, Wenjie Chen, Gaby Nahum, Junhua Gu, Peter Chu, Yucheng Liu, Xin Zhao, Vitor Cid, Chaorong Chen, Vijay Pappu, Ashwin Kumar, Wenlin Chen, Ben Schulte, Deepak Chandra, Ritwik Tewari