arXiv:2607. 26253v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) is bottlenecked by rollout generation, yet many sampled prompts produce saturated groups (all responses correct or all incorrect) whose zero reward variance yields no policy-gradient signal.
By Pixel Nomand, Elena Voss, Marcus Hale, Sofia Reyes
arXiv:2608. 07719v1 Announce Type: new Abstract: Offline reinforcement learning repeatedly trains policies from a fixed transition pool, making redundant data costly across seeds and hyperparameters, while naive subsampling can remove rare transitions needed for long-horizon credit assignment.
By Ibne Farabi Shihab, Sanjeda Akter, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb, Anuj Sharma
arXiv:2606. 14581v4 Announce Type: replace Abstract: High-throughput experimentation can evaluate many reaction conditions, yet combinatorial condition spaces still exceed the available experiment budget.
By Guanyu Liu, Weiyi Kong, Chao Tang, Zeyu Wang, Boer Zhang, Baiqing Li, Peiyu Zhang, Tianyu Shi
arXiv:2609.37932v1 Announce Type: new
Abstract: Systems operating in dynamic environments require timely updates to sustain performance. For resource-intensive systems such as machine learning models...
By Qiulin Lin, Junyan Su, Liyuan Wang, Minghua Chen
The paper introduces a request-driven framework for designing budgeted threshold incentives on on-demand delivery platforms. It decomposes the process into four stages—conditional prediction, population reduction, trajectory integration, and budget allocation—using seven interchangeable modules that share conditional trajectory laws. The framework includes a response-correction step that reweights abundant no-offer data to match short pilot moments, and the authors prove that the end-to-end value loss is bounded by the sum of stage errors, with empirical results showing significant speedups and reduced regret compared to traditional trials.
By Zhuolin Wu, Chengrui Zhu, Wenhua Nie, Kenny Ye Liang, Junming Lin, Haiyang Li, Zhilin Li, Wenjia Geng, Zeyu Wu, Yinan Wu, Jinghua Hao, Renqing He
arXiv:2608. 03447v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive generation by verifying a draft block with a target model in parallel.
By Yuannuo Feng, Zegang Peng, Yuxin Xie, Yubing Ye, Yizhe Chen, Wenshuai Yao, Wenyong Zhou, Wang Kang
arXiv:2609.39634v1 Announce Type: cross
Abstract: Common policy improvement methods, including TRPO, PPO, and GRPO, estimate policy improvement under the behavioral policy's state-visitation distribu...
By Nima H. Siboni
arXiv:2608. 11318v1 Announce Type: cross Abstract: Many sequential construction tasks exhibit exact symmetry at completion while their execution remains directed and history-dependent.
By Yi Liu
arXiv:2602. 09456v2 Announce Type: replace Abstract: We propose an algorithmic framework, Offline Estimation to Decisions (OE2D), that efficiently reduces contextual bandit learning with general reward function approximation to offline regression.
By Hao Qin, Chicheng Zhang
The paper introduces Regret-Weighted Payoff Sampling (RWPS), a budgeted estimator that selectively simulates only payoff-matrix cells relevant to a Nash equilibrium and uses a surrogate model for the remaining entries. RWPS provides an instance-dependent error bound weighted by the opponent’s equilibrium mixture and a coverage result guaranteeing that, once the deviation-relevant set is simulated, surrogate error does not affect either player’s regret. Experiments on three 21×21 general-sum games, including an asymmetric Colonel Blotto, show that RWPS achieves four to six times tighter bounds than previous methods and outperforms other sampling strategies on the CyGym and ANSG cyber simulators at low budgets.
By Michael Lanier, David Farmer, Yevgeniy Vorobeychik
The paper introduces loss‑conditioned state execution, a model‑agnostic technique that decides whether to apply a world model’s proposed state change or keep the current state based on whether the change reduces downstream loss. It formalizes state movability as the existence of a loss‑reducing feasible correction and constructs loss‑specific proposals from predictive distributions, executing them only when a groupwise lower confidence bound on loss improvement is positive. Experiments on forecasting and dynamics benchmarks show that the method accepts updates for a subset of cases, achieving lower bounded loss than persistence or always executing the proposal, and highlights that event predictability and loss‑based decisions must be evaluated separately.
By Jintao Xu, Zhengyu Chen, Ben Zhang, Yongzhi Qi, Jianshen Zhang
arXiv:2607. 02255v1 Announce Type: new Abstract: Memory for a long-horizon LLM agent is a contract about what each future decision is allowed to see.
By Xiangchen Cheng, Yunwei Jiang, Jianwen Sun, Zizhen Li, Chuanhao Li, Xiangcheng Cao, Yihao Liu, Fanrui Zhang, Li Jin, Kaipeng Zhang