arXiv AI By Weiliang Zhang, Xiaohan Huang, Yi Du, Ziyue Qiao, Qingqing Long, Zhen Meng, Yuanchun Zhou, Meng Xiao

Comprehend, Divide, and Conquer: Feature Subspace Exploration via Multi-Agent Hierarchical Reinforcement Learning

Read the original on arXiv AI →

arXiv:2504. 17356v3 Announce Type: replace Abstract: Feature selection aims to preprocess the target dataset, find an optimal and most streamlined feature subset, and enhance the downstream machine learning task.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 12

Discovering Temporal Structure: An Overview of Hierarchical Reinforcement Learning

The paper reviews hierarchical reinforcement learning (HRL) as a method for enabling agents to explore, plan, and learn in complex, open-ended environments by uncovering temporal structure in experience streams. It discusses the unclear definition of what makes a structure useful, the benefits of HRL for decision‑making challenges, and its impact on AI agent performance trade‑offs. The authors survey various HRL methods—from online learning to offline datasets and large language model integration—and outline the challenges and suitable domains for temporal structure discovery.

By Martin Klissarov, Akhil Bagaria, Ziyan Luo, George Konidaris, Doina Precup, Marlos C. Machado
arXiv AI
Sep 21

Collab-Solver: Collaborative Solving Policy Learning for Mixed-Integer Linear Programming

Collab‑Solver introduces a multi‑agent policy learning framework for mixed‑integer linear programming (MILP) that enables collaborative optimization of multiple solver modules. By modeling the interaction between cut selection and branching as a Stackelberg game, the approach employs a two‑phase learning paradigm—data‑communicated policy pretraining followed by coordinated policy refinement. Experiments on synthetic and large‑scale real‑world MILP datasets show that the jointly learned policies markedly improve solving performance and generalize well across diverse instance sets.

By Siyuan Li, Yifan Yu, Zhihao Zhang, Mengjing Chen, Fangzhou Zhu, Tao Zhong, Peng Liu, Jianye Hao