arXiv AI By Zhuolin Wu, Chengrui Zhu, Wenhua Nie, Kenny Ye Liang, Junming Lin, Haiyang Li, Zhilin Li, Wenjia Geng, Zeyu Wu, Yinan Wu, Jinghua Hao, Renqing He

A General Framework for Budgeted Threshold Incentives on Request

Read the original on arXiv AI →

The paper introduces a request-driven framework for designing budgeted threshold incentives on on-demand delivery platforms. It decomposes the process into four stages—conditional prediction, population reduction, trajectory integration, and budget allocation—using seven interchangeable modules that share conditional trajectory laws. The framework includes a response-correction step that reweights abundant no-offer data to match short pilot moments, and the authors prove that the end-to-end value loss is bounded by the sum of stage errors, with empirical results showing significant speedups and reduced regret compared to traditional trials.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 30

Early Verdicts, Better Budgets: Sequential Adaptive Rollout Allocation for Compute-Efficient RLVR

arXiv:2607. 26253v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by rollout generation, yet many sampled prompts produce saturated groups (all responses correct or all incorrect) whose zero reward variance yields no policy-gradient signal.

By Pixel Nomand, Elena Voss, Marcus Hale, Sofia Reyes
arXiv AI
6d ago

ERRAND: Budgeted Maintenance of Agent Memory

ERRAND is a new method for budgeted maintenance of agent memory that treats revalidation of stored knowledge as a priced errand competing for scarce actions. It uses an errand index that is single‑peaked, allowing certainty in either direction to cost nothing, and repairs by writing new versions rather than deleting old ones. In experiments across two drifting tool‑use worlds, ERRAND outperforms non‑oracle policies, achieving up to 10.0 percentage points improvement over eager revalidation while using only 11.0% of steps, and it self‑terminates when no budget is imposed.

By Beining Wu, Zihao Ding, Jun Huang