arXiv AI

Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation

arXiv:2609. 29140v1 Announce Type: new Abstract: Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty.

arXiv Machine Learning
Jul 30

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR

arXiv:2607. 26253v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by rollout generation, yet many sampled prompts produce saturated groups (all responses correct or all incorrect) whose zero reward variance yields no policy-gradient signal.

By Pixel Nomand, Elena Voss, Marcus Hale, Sofia Reyes
arXiv AI
6d ago

Certified Task-Conditioned Active Observability

The paper introduces the concept of task‑conditioned active observability, defining the minimal interaction cost needed for an autonomous agent to identify task‑relevant states while guaranteeing safe abstention. It formalizes this complexity, proving that task‑predictive equivalence yields a unique minimal sufficient quotient that preserves complexity and eliminates unnecessary distinctions. The authors present theoretical characterizations for deterministic and noisy regimes, and demonstrate a certified observer that reduces sensor usage and model steps while maintaining zero false acceptances in extensive high‑dimensional trials.

By Linzhe Zhang, Changming Xu
arXiv Machine Learning
1d ago

Tail-Influence Sampling for CVaR Policy Evaluation

arXiv:2609.38096v1 Announce Type: new Abstract: Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can requi...

By Pauline Bourigault, Xiaotong Ji, Matthieu Zimmer, Rasul Tutunov, Haitham Bou-Ammar
arXiv AI
Sep 10

Beyond "AI Helps Humans": Decision-Targeted Evaluation Design for Human-Agent Teams in the Agentic Era

The paper introduces TEAM-Design, a rule that assigns two replay probabilities to each task—one for a human-only replay and one for an agent-only replay—based on how difficult it is to predict the missing baseline outcome and the cost of replay. It addresses the challenge of deciding whether to keep a human-AI workflow or replace it with a single actor when only one outcome can be observed after deployment. The authors prove that TEAM-Design solves the budgeted design problem and controls error rates, and demonstrate its effectiveness on clinical and coding benchmarks, noting it excels when one comparison is clearly harder than the other.

By Hamed Khosravi, Xiaoming Huo