arXiv AI

Comprehend, Divide, and Conquer: Feature Subspace Exploration via Multi-Agent Hierarchical Reinforcement Learning

arXiv:2504. 17356v3 Announce Type: replace Abstract: Feature selection aims to preprocess the target dataset, find an optimal and most streamlined feature subset, and enhance the downstream machine learning task.

arXiv AI
Sep 12

Discovering Temporal Structure: An Overview of Hierarchical Reinforcement Learning

The paper reviews hierarchical reinforcement learning (HRL) as a method for enabling agents to explore, plan, and learn in complex, open-ended environments by uncovering temporal structure in experience streams. It discusses the unclear definition of what makes a structure useful, the benefits of HRL for decision‑making challenges, and its impact on AI agent performance trade‑offs. The authors survey various HRL methods—from online learning to offline datasets and large language model integration—and outline the challenges and suitable domains for temporal structure discovery.

By Martin Klissarov, Akhil Bagaria, Ziyan Luo, George Konidaris, Doina Precup, Marlos C. Machado
arXiv AI
Sep 21

Collab-Solver: Collaborative Solving Policy Learning for Mixed-Integer Linear Programming

Collab‑Solver introduces a multi‑agent policy learning framework for mixed‑integer linear programming (MILP) that enables collaborative optimization of multiple solver modules. By modeling the interaction between cut selection and branching as a Stackelberg game, the approach employs a two‑phase learning paradigm—data‑communicated policy pretraining followed by coordinated policy refinement. Experiments on synthetic and large‑scale real‑world MILP datasets show that the jointly learned policies markedly improve solving performance and generalize well across diverse instance sets.

By Siyuan Li, Yifan Yu, Zhihao Zhang, Mengjing Chen, Fangzhou Zhu, Tao Zhong, Peng Liu, Jianye Hao
arXiv Machine Learning
Aug 27

Regularized Latent Dynamics Prediction is a Strong Baseline For Behavioral Foundation Models

The paper introduces Regularized Latent Dynamics Prediction (RLDP), a method that adds orthogonality regularization to self‑supervised next‑state prediction in latent space. RLDP maintains feature diversity, matching or surpassing complex representation learning approaches for zero‑shot reinforcement learning. It also performs robustly in low‑coverage data settings where prior methods fail.

By Pranaya Jajoo, Harshit Sikchi, Siddhant Agarwal, Amy Zhang, Scott Niekum, Martha White
arXiv AI
Jun 10

HIPIF: Hierarchical Planning and Information Folding for Long-Horizon LLM Agent Learning

arXiv:2606. 10507v1 Announce Type: new Abstract: While Large Language Models (LLMs) have demonstrated strong capabilities as autonomous agents across a wide range of tasks, their performance often degrades in multi-turn long-horizon agentic tasks.

By Juncheng Diao, Zhicong Lu, Peiguang Li, Yongwei Zhou, Changyuan Tian, Qingbin Li, Rongxiang Weng, Jingang Wang, Xunliang Cai