arXiv Machine Learning By Silviu Pitis

Rationalizing Boltzmann Rationality: An Axiomatic Characterization of Entropy-Regularized Policies

Read the original on arXiv Machine Learning →

arXiv:2607. 17316v1 Announce Type: new Abstract: The softmax policy $\pi(a \mid s) \propto \exp(\beta Q(s,a))$ is the default model of stochastic choice in reinforcement learning (RL).

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
5d ago

Annealed Softmax Greedy in Many-Armed Bayesian Bandits

arXiv:2605. 31034v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) and group-based policy optimization methods such as GRPO update a stochastic policy by sampling multiple completions per prompt and increasing the policy's probability on those with higher reward, regularized by a KL penalty toward a reference policy.

By William Overman, Mohsen Bayati