The paper introduces a framework for hypothesis testing that combines inexpensive AI judgments with selective human verification to control type‑I and type‑II errors while minimizing cost. It derives an information‑theoretic lower bound on the minimum cost and proposes the SCALE policy, a sequential, cost‑aware strategy that adapts AI scoring and human escalation. SCALE is proven valid for finite samples and asymptotically matches the lower bound, achieving significant savings when both AI and human inputs are valuable.
By Dae Woong (David), Ham, Xuejun Zhao, Stefanus Jasin, Fenghua Yang
arXiv:2602. 08335v2 Announce Type: replace Abstract: Integrating Large Language Models (LLMs) with external tools via multi-agent systems offers a promising new paradigm for decomposing and solving complex problems.
By Yanming Li, Xuelin Zhang, WenJie Lu, Ziye Tang, Maodong Wu, Haotian Luo, Tongtong Wu, Zijie Peng, Hongze Mi, Yibo Feng, Naiqiang Tan, Chao Huang, Lian Peng, Li Shen
arXiv:2507. 09473v2 Announce Type: replace-cross Abstract: We study the dynamic allocation of indivisible resources to strategic agents under long-term constraints, where the planner aims to maximize social welfare, satisfy multiple constraints, and elicit near-truthful reports.
By Yan Dai, Negin Golrezaei, Patrick Jaillet
arXiv:2603. 17212v2 Announce Type: replace-cross Abstract: When organizations delegate text generation tasks to AI providers via pay-for-performance contracts, expected payments rise when evaluation is noisy.
By Eden Saig, Tamar Garbuz, Ariel D. Procaccia, Inbal Talgam-Cohen, Jamie Tucker-Foltz
The paper proposes a method for allocating a limited budget of expert annotations to optimize the accuracy of off-policy evaluation in settings where rewards are missing or noisy. By deriving variance‑optimal annotation probabilities for sequential, forward‑monotone protocols, the authors provide a batch‑adaptive implementation that can be applied to real data. Experiments on casenotes from a homelessness services nonprofit and on human‑preference votes from LMArena demonstrate substantial reductions in RMSE—up to 65% for housing placement and 68% for progress toward a housing application—when using only 40% or more of the full annotation budget.
By Woojin Chae, Ezinne Nwankwo, Haitong Qin, Angela Zhou
arXiv:2602.04125v2 Announce Type: replace-cross
Abstract: Modern digital platforms use contextual bandits to allocate valuable exposure and opportunities among competing participants. Fair treatment...
By Qingwen Zhang, Wenjia Wang