arXiv Machine Learning

Smooth Learning with Hard Constraints via Legendre-Regularized Policies

arXiv:2607. 24007v1 Announce Type: cross Abstract: We revisit contextual optimization from the perspective of policy class design.

arXiv Machine Learning
Jun 9

Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial Actions

arXiv:2601. 22211v2 Announce Type: replace Abstract: Reinforcement learning (RL) with combinatorial action spaces remains challenging because feasible action sets are exponentially large and governed by complex feasibility constraints, making direct policy parameterization impractical.

By Lingkai Kong, Anagha Satish, Hezi Jiang, Akseli Kangaslahti, Andrew Ma, Wenbo Chen, Mingxiao Song, Lily Xu, Milind Tambe
arXiv AI
Jun 6

Retry Policy Gradients in Continuous Action Spaces

arXiv:2606. 05888v1 Announce Type: new Abstract: Retry-based objectives such as pass@K and max@K optimize the best return obtained from multiple sampled trajectories, and recent work has shown that they can promote exploration without explicit exploration bonuses.

By Soichiro Nishimori, Paavo Parmas