arXiv Machine Learning

Mastering Atari 2600 Games with Discovered Options

Wayfarer is a domain‑agnostic, online deep RL agent that discovers options via Laplacian representation learning from high‑dimensional observations and uses them for control. The discovered options improve exploration, accelerate credit assignment, and generalise to unseen settings, leading to faster learning of complex policies. Wayfarer achieves state‑of‑the‑art performance among single‑stream agents on the most challenging Atari 2600 games, especially those requiring long‑horizon exploration such as Montezuma's Revenge and Private Eye.

arXiv AI
Sep 3

Action abstractions for amortized sampling

The paper introduces a method that integrates action abstraction into policy optimization for reinforcement learning and generative flow networks. By iteratively identifying frequently used action subsequences in high‑reward trajectories and treating them as single high‑level actions, the approach expands the action space and improves sample efficiency. Experiments on synthetic and real‑world tasks show that this technique discovers diverse high‑reward states more effectively, especially on challenging exploration problems, and yields interpretable abstract actions that reflect the underlying reward structure.

By Oussama Boussif, L\'ena N\'ehale Ezzine, Joseph D Viviano, Micha{\l} Koziarski, Moksh Jain, Esmeralda S. Whitammer, Emmanuel Bengio, Rim Assouel, Yoshua Bengio
arXiv Machine Learning
Oct 2

Towards More Efficient, Robust, Instance-adaptive, and Generalizable Sequential Decision making

The article outlines a Ph.D. research agenda aimed at creating provably efficient and practical algorithms for data‑driven sequential decision‑making under uncertainty. It focuses on reinforcement learning and multi‑armed bandits, targeting applications such as recommendation systems, computer networks, video analytics, and large language models. The work seeks to overcome limitations of existing methods—such as reliance on idealized models, lack of robustness to adversarial perturbations, and poor instance‑dependent performance—by developing algorithms that are more efficient, robust, instance‑adaptive, and generalizable to new environments.

By Zhiyong Wang
arXiv AI
Jul 23

In-the-Flow Agentic System Optimization for Effective Planning and Tool Use

arXiv:2510. 05592v2 Announce Type: replace Abstract: Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interleaves thoughts and tool calls under full context; this scales poorly with long horizons and diverse tools and generalizes weakly to new scenarios.

By Zhuofeng Li, Haoxiang Zhang, Seungju Han, Sheng Liu, Jianwen Xie, Yu Zhang, Yejin Choi, James Zou, Pan Lu
arXiv Computation and Language
Sep 21

ArenaFlow: From Trajectory Ranking to Hierarchical Credit Propagation for Open-Ended Agent RL

ArenaFlow is a hierarchical credit propagation framework designed to improve reinforcement learning for open-ended agent tasks. It uses tournament-based relative ranking to generate trajectory-level rewards and structured reflective evaluation to identify pivotal success steps, reusable strategy skills, and skill usage attribution. The framework propagates advantages to high-confidence steps and maintains a global skill memory, enabling more targeted optimization and reusable skill priors for future exploration.

By Qiang Zhang, Ruixue Ding, Fanrui Zhang, Xi Chen, Boli Chen, Shihang Wang, Yinfeng Huang, Yi Zheng, Pengjun Xie, Kaipeng Zhang, Jiawei Liu, Zheng-Jun Zha