arXiv:2609.36740v1 Announce Type: new
Abstract: Many recommender systems such as for e-commerce and news platforms aim to provide users with rankings they are likely to interact with. Off-Policy Lear...
By Ren Kishimoto, Koichi Tanaka, Haruka Kiyohara, Yusuke Narita, Yasuo Yamamoto, Nobuyuki Shimizu, Yuta Saito
The paper introduces a new algorithm for the Multi‑Armed Bandit problem that prioritizes selecting the arm with the lowest variance rather than the highest expected reward, using a softmax policy parameterization. It constructs an unbiased estimate of the minimal‑variance objective by drawing two independent samples from the chosen arm and proves convergence under natural conditions. Numerical experiments demonstrate the algorithm’s practical behavior and provide implementation guidance, while also addressing general risk‑aware trade‑offs between average reward and variance.
By Gabriel Turinici
arXiv:2409. 03682v2 Announce Type: replace Abstract: Learning new tasks by leveraging prior experience is a fundamental trait of intelligent systems.
By El Mahdi Chayti, Martin Jaggi
arXiv:2607. 11146v1 Announce Type: new Abstract: We study the coupled objective J_K^WOR = E_{S ~ PL-WOR_K}[max_{i in S} R_i]: the expected maximum reward of a size-K Plackett-Luce draw without replacement, the law of Gumbel-Top-K / Stochastic Beam Search decoding.
By Melveena Jolly, Midhun Xavier
arXiv:2505. 15201v5 Announce Type: replace-cross Abstract: Reinforcement Learning (RL) algorithms sample multiple n>1 solution attempts for each problem and reward them independently.
By Christian Walder, Deep Karkhanis
arXiv:2606. 06080v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards is widely used for post-training reasoning models, but sparse outcome rewards make exploration difficult.
By Shota Takashiro, Soichiro Nishimori, Paavo Parmas, Yongmin Kim, Kohsei Matsutani, Gouki Minegishi, Yusuke Iwasawa, Takeshi Kojima, Yutaka Matsuo
The paper introduces Adaptive Doubly Robust (ADR), an off‑policy evaluation method for ranking policies that blends adaptive importance weighting with reward regression to reduce variance. ADR is unbiased when the true user behavior model is known and, under a sufficient condition, achieves lower variance than the prior Adaptive Inverse Propensity Scoring (AIPS) approach. Experiments on synthetic data show that ADR consistently improves mean squared error over AIPS and other ranking OPE estimators across various data sizes and ranking lengths.
By Kosuke Iguchi, Ren Kishimoto
arXiv:2606. 01081v1 Announce Type: new Abstract: Decision-focused learning (DFL) trains predictive models by optimizing downstream decision quality rather than standalone prediction accuracy.
By Wyame Benslimane, Tinghan Ye, Pascal Van Hentenryck, Paul Grigas
arXiv:2607. 00691v1 Announce Type: new Abstract: Black-box optimization is a fundamental science and engineering tool that makes it possible to optimize objectives without gradient information.
By Edouard R. Dufour, Pascal Fua
arXiv:2606. 08480v1 Announce Type: cross Abstract: Reinforcement learning (RL) presents a promising avenue for enhancing generative recommendation beyond supervised imitation, leveraging reward signals to guide policy improvement.
By Kewei Xu, Junbo Qi, Yanyan Zou, Pengfei Zhang, Xingzhi Yao, Shengjie Li
arXiv:2509. 22047v3 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) has been shown to be an effective algorithm when an accurate reward model is available.
By Yuki Ichihara, Yuu Jinnai, Tetsuro Morimura, Mitsuki Sakamoto, Ryota Mitsuhashi, Eiji Uchibe
arXiv:2606. 03831v1 Announce Type: new Abstract: This paper investigates non-stationary online learning using the metric of interval regret, which requires an online algorithm to perform well over every time interval.
By Yan-Feng Xie, Shuche Wang, Peng Zhao, Zhi-Hua Zhou