arXiv:2608. 14761v1 Announce Type: cross Abstract: At a finite public-chance cut, counterfactual regret minimization (CFR) must choose how many outcomes to evaluate before each regret update.
By Jiaxing Guo, Lei Ye
arXiv:2609.38914v1 Announce Type: new
Abstract: Evaluating interactive agents is expensive. Agent behavior is stochastic, so reliability must be measured over repeated trials, but failures are rare a...
By Priyanath Maji, Spandan Ghose Chowdhury
arXiv:2607. 26253v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by rollout generation, yet many sampled prompts produce saturated groups (all responses correct or all incorrect) whose zero reward variance yields no policy-gradient signal.
By Pixel Nomand, Elena Voss, Marcus Hale, Sofia Reyes
ERRAND is a new method for budgeted maintenance of agent memory that treats revalidation of stored knowledge as a priced errand competing for scarce actions. It uses an errand index that is single‑peaked, allowing certainty in either direction to cost nothing, and repairs by writing new versions rather than deleting old ones. In experiments across two drifting tool‑use worlds, ERRAND outperforms non‑oracle policies, achieving up to 10.0 percentage points improvement over eager revalidation while using only 11.0% of steps, and it self‑terminates when no budget is imposed.
By Beining Wu, Zihao Ding, Jun Huang
arXiv:2609. 29140v1 Announce Type: new Abstract: Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty.
By Yezhou Cheng, Runjia Du, Zeming Liu, Qibai Chen, Hang Lyu, Yilan Wei, Yankai Zeng, Bojun Lin
arXiv:2608. 13209v1 Announce Type: cross Abstract: Many operational decisions are sequences of interventions under a cumulative resource limit, such as a maintenance schedule within a crew-hour budget.
By Minkyoung Kim, Beakcheol Jang