arXiv Machine Learning By Paavo Parmas, Yongmin Kim, Kohsei Matsutani, Shota Takashiro, Soichiro Nishimori, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation

Read the original on arXiv Machine Learning →

arXiv:2606. 06096v1 Announce Type: new Abstract: Policy-gradient methods usually optimize expected return, but many real world applications care about distributional properties of returns: tail risk, outlier robustness, or best-of-K discovery.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 24

Softmax gradient policy for variance minimization and risk-averse multi armed bandits

The paper introduces a new algorithm for the Multi‑Armed Bandit problem that prioritizes selecting the arm with the lowest variance rather than the highest expected reward, using a softmax policy parameterization. It constructs an unbiased estimate of the minimal‑variance objective by drawing two independent samples from the chosen arm and proves convergence under natural conditions. Numerical experiments demonstrate the algorithm’s practical behavior and provide implementation guidance, while also addressing general risk‑aware trade‑offs between average reward and variance.

By Gabriel Turinici
arXiv Machine Learning
Jun 5

On Advantage Estimates for Max@K Policy Gradients

arXiv:2606. 06080v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards is widely used for post-training reasoning models, but sparse outcome rewards make exploration difficult.

By Shota Takashiro, Soichiro Nishimori, Paavo Parmas, Yongmin Kim, Kohsei Matsutani, Gouki Minegishi, Yusuke Iwasawa, Takeshi Kojima, Yutaka Matsuo