arXiv:2607. 23854v1 Announce Type: new Abstract: Humans often find good solutions to combinatorial optimization problems that are computationally hard even for advanced computer algorithms.
By Haijiang Yan, Jian-Qiao Zhu, Liqiang Huang, Ming Meng
Finding policies for Markov Decision Processes (MDPs) is a central problem in areas such as Reinforcement Learning and Operations Research. Here, we have to repeatedly choose an action that should be...
arXiv:2405. 13947v2 Announce Type: replace Abstract: Deep neural networks based on reinforcement learning (RL) for solving combinatorial optimization (CO) problems are developing rapidly and have shown a tendency to approach or even outperform traditional solvers.
By Chaoyang Wang, Pengzhi Cheng, Jingze Li, Weiwei Sun
arXiv:2609.15539v1 Announce Type: cross
Abstract: Finding policies for Markov Decision Processes (MDPs) is a central problem in areas such as Reinforcement Learning and Operations Research. Here, we...
By Lars Rohwedder, Rico Zenklusen
arXiv:2607. 06066v1 Announce Type: new Abstract: The Vehicle Routing Problem (VRP) and its variants represent some of the most practically consequential optimization challenges in modern logistics and urban mobility.
By Manish Kolachalam, Rani Malhotra
arXiv:2509. 18930v3 Announce Type: replace-cross Abstract: Neural algorithmic reasoning (NAR) is a paradigm that trains neural networks to execute classic algorithms by supervised learning.
By Alex Schutz, Victor-Alexandru Darvariu, Efimia Panagiotaki, Bruno Lacerda, Nick Hawes
arXiv:2606. 04860v1 Announce Type: cross Abstract: Finding optimal solution paths for combinatorial puzzles like the Rubik's Cube, sliding tile puzzles, and Lights Out remains a classical challenge in artificial intelligence.
By Siddharth Sahay
arXiv:2601. 13465v4 Announce Type: replace Abstract: Graph neural networks are usually treated as auxiliaries for combinatorial optimization: they imitate algorithms, guide search, or supply scores to classical procedures.
By Yimeng Min, Carla P. Gomes
The paper introduces the Orienteering Problem with Uncertain Time‑Varying Rewards (OP‑UTVR), a new variant of the classic orienteering problem that allows agents to estimate and forecast reward dynamics from observations. Three planners with different planning horizons and online adaptivity are proposed, and theoretical performance bounds under reward stochasticity are derived. A mobile service robot benchmark is presented, and experiments show trade‑offs between planning horizon and adaptivity, highlighting the benefits of long‑horizon planning with online adaptation.
By Masafumi Endo, Kohei Honda, Yuu Jinnai, Ryo Yonetani
arXiv:2606. 04167v1 Announce Type: cross Abstract: We tackle the Metro Network Expansion Problem (MNEP), a subset of the Transport Network Design Problem (TNDP), which focuses on expanding metro systems to satisfy travel demand.
By Dimitris Michailidis, Sennay Ghebreab, Fernando P. Santos
arXiv:2606. 01425v1 Announce Type: new Abstract: Mixed-combinatorial nonlinear programming (MCNLP) problems arise in many engineering design and planning applications, e.
By Gishnu Madhu, Feng Liu, Souma Chowdhury
The paper presents an end‑to‑end framework that uses constraint‑oriented hypergraphs and reinforcement learning to solve vehicle routing problems. It introduces a dynamic hyperedge reconstruction strategy for better hypergraph representation and a double‑pointer attention decoder for iterative solution generation. Experiments on benchmark datasets show that the method removes the need for complex heuristic operators while improving solution quality.
By Zhenwei Wang, Tiehua Zhang, Jing Liu, Heng Yu, Kaizhu Huang, Ruibin Bai