arXiv:2609.36473v1 Announce Type: new
Abstract: Temporal abstraction via options can improve exploration in large environments. However, existing option discovery algorithms find subgoals that target...
By Akhil Bagaria, Anita De Mello Koch, George Konidaris
The paper introduces a method that integrates action abstraction into policy optimization for reinforcement learning and generative flow networks. By iteratively identifying frequently used action subsequences in high‑reward trajectories and treating them as single high‑level actions, the approach expands the action space and improves sample efficiency. Experiments on synthetic and real‑world tasks show that this technique discovers diverse high‑reward states more effectively, especially on challenging exploration problems, and yields interpretable abstract actions that reflect the underlying reward structure.
By Oussama Boussif, L\'ena N\'ehale Ezzine, Joseph D Viviano, Micha{\l} Koziarski, Moksh Jain, Esmeralda S. Whitammer, Emmanuel Bengio, Rim Assouel, Yoshua Bengio
arXiv:2410.14606v3 Announce Type: replace
Abstract: Learning from a stream of experience as it arrives, also known as streaming learning, is a core part of natural learning. However, reliable streami...
By Mohamed Elsayed, Elena Sorina Lupu, Gautham Vasan, A. Rupam Mahmood
arXiv:2601. 19612v3 Announce Type: replace-cross Abstract: Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.
By Manuel Wendl, Yarden As, Manish Prajapat, Anton Pollak, Stelian Coros, Andreas Krause
arXiv:2605. 29032v2 Announce Type: replace Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss.
By Christoph Dann, Yishay Mansour, Mehryar Mohri
arXiv:2607. 00392v1 Announce Type: cross Abstract: Unsupervised Reinforcement Learning (URL) aims to pre-train scalable, skill-conditioned policies without extrinsic rewards, serving as a foundation for downstream control tasks.
By Jongchan Park, Seungjun Oh, Seungho Baek, Yusung Kim
The article outlines a Ph.D. research agenda aimed at creating provably efficient and practical algorithms for data‑driven sequential decision‑making under uncertainty. It focuses on reinforcement learning and multi‑armed bandits, targeting applications such as recommendation systems, computer networks, video analytics, and large language models. The work seeks to overcome limitations of existing methods—such as reliance on idealized models, lack of robustness to adversarial perturbations, and poor instance‑dependent performance—by developing algorithms that are more efficient, robust, instance‑adaptive, and generalizable to new environments.
By Zhiyong Wang
arXiv:2510. 05592v2 Announce Type: replace Abstract: Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interleaves thoughts and tool calls under full context; this scales poorly with long horizons and diverse tools and generalizes weakly to new scenarios.
By Zhuofeng Li, Haoxiang Zhang, Seungju Han, Sheng Liu, Jianwen Xie, Yu Zhang, Yejin Choi, James Zou, Pan Lu
arXiv:2602. 05999v3 Announce Type: replace Abstract: How does the amount of compute available to a reinforcement learning (RL) policy affect its learning?
By Raj Ghugare, Micha{\l} Bortkiewicz, Alicja Ziarko, Benjamin Eysenbach
arXiv:2508. 14751v2 Announce Type: replace Abstract: We study goal-conditioned reinforcement learning in partially observable environments with sparse rewards and large, structured goal spaces.
By Thomas Carta, Cl\'ement Romac, Loris Gaven, Pierre-Yves Oudeyer, Olivier Sigaud, Sylvain Lamprier
ArenaFlow is a hierarchical credit propagation framework designed to improve reinforcement learning for open-ended agent tasks. It uses tournament-based relative ranking to generate trajectory-level rewards and structured reflective evaluation to identify pivotal success steps, reusable strategy skills, and skill usage attribution. The framework propagates advantages to high-confidence steps and maintains a global skill memory, enabling more targeted optimization and reusable skill priors for future exploration.
By Qiang Zhang, Ruixue Ding, Fanrui Zhang, Xi Chen, Boli Chen, Shihang Wang, Yinfeng Huang, Yi Zheng, Pengjun Xie, Kaipeng Zhang, Jiawei Liu, Zheng-Jun Zha
Deliberating takes time. In real-time settings, that time is not free.