A Framework for Measuring Appropriate Reliance on Set-Valued AI Advice
arXiv:2606. 06081v1 Announce Type: new Abstract: Appropriate reliance on AI advice has become a central research theme in human-AI collaboration.
arXiv:2606. 12587v1 Announce Type: new Abstract: Traditionally, decision support studies how humans use machine learning models to make better decisions.
arXiv:2606. 06081v1 Announce Type: new Abstract: Appropriate reliance on AI advice has become a central research theme in human-AI collaboration.
arXiv:2608. 06714v1 Announce Type: new Abstract: Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods.
Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods. We ask a fundamentally different question: how much of this search policy can be internalized by a single tool-using agent?
The paper introduces a framework for hypothesis testing that combines inexpensive AI judgments with selective human verification to control type‑I and type‑II errors while minimizing cost. It derives an information‑theoretic lower bound on the minimum cost and proposes the SCALE policy, a sequential, cost‑aware strategy that adapts AI scoring and human escalation. SCALE is proven valid for finite samples and asymptotically matches the lower bound, achieving significant savings when both AI and human inputs are valuable.
arXiv:2606. 13468v1 Announce Type: cross Abstract: AI coding agents are increasingly used to generate pull requests (PRs) that propose code fixes in software projects.
The paper introduces TEAM-Design, a rule that assigns two replay probabilities to each task—one for a human-only replay and one for an agent-only replay—based on how difficult it is to predict the missing baseline outcome and the cost of replay. It addresses the challenge of deciding whether to keep a human-AI workflow or replace it with a single actor when only one outcome can be observed after deployment. The authors prove that TEAM-Design solves the budgeted design problem and controls error rates, and demonstrate its effectiveness on clinical and coding benchmarks, noting it excels when one comparison is clearly harder than the other.
arXiv:2607. 05775v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and operate over extended horizons.
The paper argues that large language models need adaptive reasoning rather than fixed reasoning budgets. It shows that over‑reasoning leads to high computational cost without accuracy gains, while under‑reasoning results in incorrect or incomplete solutions. The authors evaluate these failure modes on MATH‑500 and the GAIA benchmark, highlighting the need for dynamic reasoning allocation in agentic AI systems.
arXiv:2403. 16178v2 Announce Type: replace-cross Abstract: For effective human-agent teaming, robots and other artificial intelligence (AI) agents must infer their human partner's abilities and behavioral response patterns and adapt accordingly.
arXiv:2607. 21268v1 Announce Type: cross Abstract: In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable correctness signal exists.
The paper introduces ProSE, a framework for AI assistants that generate proposals while considering users’ bounded rationality and evaluability constraints. It proposes a KL‑regularised bounded‑rational binary response model and a depth‑2 Bayes‑adaptive planner, “ProSE‑Plan,” which scores proposals by expected responses and resulting belief updates. Experiments on graph simulations show that “ProSE‑Plan” outperforms evaluability‑unaware and myopic baselines, especially when evaluation cost is high, and that informative probes are crucial for effective assistance.
arXiv:2609.08248v1 Announce Type: new Abstract: Modern industrial ads ranking stacks are increasingly bottlenecked not by model capacity or training compute, but by the throughput of human ML iterati...