arXiv:2606. 06081v1 Announce Type: new Abstract: Appropriate reliance on AI advice has become a central research theme in human-AI collaboration.
By Ranjan Mishra, Jakob Schoeffer
arXiv:2608. 06714v1 Announce Type: new Abstract: Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods.
By Junbo Li, Boyi Liu, Canwen Xu, Yite Wang, Yuxiong He, Zhangyang Wang, Qiang Liu, Zhewei Yao
Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods. We ask a fundamentally different question: how much of this search policy can be internalized by a single tool-using agent?
The paper introduces a framework for hypothesis testing that combines inexpensive AI judgments with selective human verification to control type‑I and type‑II errors while minimizing cost. It derives an information‑theoretic lower bound on the minimum cost and proposes the SCALE policy, a sequential, cost‑aware strategy that adapts AI scoring and human escalation. SCALE is proven valid for finite samples and asymptotically matches the lower bound, achieving significant savings when both AI and human inputs are valuable.
By Dae Woong (David), Ham, Xuejun Zhao, Stefanus Jasin, Fenghua Yang
arXiv:2606. 13468v1 Announce Type: cross Abstract: AI coding agents are increasingly used to generate pull requests (PRs) that propose code fixes in software projects.
By Mahmoud Abujadallah, Ali Arabat, Mohammed Sayagh
The paper introduces TEAM-Design, a rule that assigns two replay probabilities to each task—one for a human-only replay and one for an agent-only replay—based on how difficult it is to predict the missing baseline outcome and the cost of replay. It addresses the challenge of deciding whether to keep a human-AI workflow or replace it with a single actor when only one outcome can be observed after deployment. The authors prove that TEAM-Design solves the budgeted design problem and controls error rates, and demonstrate its effectiveness on clinical and coding benchmarks, noting it excels when one comparison is clearly harder than the other.
By Hamed Khosravi, Xiaoming Huo