arXiv:2607. 23854v1 Announce Type: new Abstract: Humans often find good solutions to combinatorial optimization problems that are computationally hard even for advanced computer algorithms.
By Haijiang Yan, Jian-Qiao Zhu, Liqiang Huang, Ming Meng
arXiv:2509.05084v2 Announce Type: replace
Abstract: The primary paradigm in Neural Combinatorial Optimization (NCO) consists of construction methods, where a neural network is trained to sequentially...
By Tim Dernedde, Daniela Thyssens, Lars Schmidt-Thieme
The paper investigates how attention dynamics evolve across recurrent depth in language models, finding that attention support stabilizes early while hidden states and outputs take longer. It proposes WISE, a training‑free method that uses full attention in early steps and then reuses the discovered sparse working set for later steps, preserving performance on multi‑hop QA tasks. Experiments show that WISE maintains quality up to 2K context, offers measurable speedups, and highlights the importance of recurrent discovery of attention support.
By Ke Wan, Chen Chen
arXiv:2601. 13465v4 Announce Type: replace Abstract: Graph neural networks are usually treated as auxiliaries for combinatorial optimization: they imitate algorithms, guide search, or supply scores to classical procedures.
By Yimeng Min, Carla P. Gomes
arXiv:2608. 07335v1 Announce Type: cross Abstract: Recent advancements in deep reinforcement learning have increasingly favored simplified, highly parallelized paradigms.
By Taha Shieenavaz, Shabnam Zareshahraki, Loris Nanni
arXiv:2506. 03672v2 Announce Type: replace-cross Abstract: Combinatorial Optimization problems are widespread in domains such as logistics, manufacturing, and drug discovery, yet their NP-hard nature makes them computationally challenging.
By Sobihan Surendran (LPSM), Adeline Fermanian (LPSM), Sylvain Le Corff (LPSM)
arXiv:2512. 14617v2 Announce Type: replace-cross Abstract: Many practical decision-making problems involve tasks whose success depends on the entire system history, rather than on achieving a state with desired properties.
By Alessandro Trapasso, Luca Iocchi, Fabio Patrizi
MiLoop is a reinforcement‑learning‑based constructive framework for neural combinatorial optimization that propagates selective memory across rollout steps. By fusing current embeddings with historical memory before attention layers and applying adaptive gated updates afterward, it enables a shallow policy to learn dynamic embeddings without external solution labels or search‑space pruning. Experiments on four combinatorial optimization problems show MiLoop consistently generates high‑quality solutions for instances ranging from 100 to 10 million nodes, demonstrating strong generalization.
By Changliang Zhou, Yuanyao Chen, Rongsheng Chen, Zhiyun Lin, Zhenkun Wang
arXiv:2607. 08340v1 Announce Type: cross Abstract: Q-learning is a fundamental algorithm in reinforcement learning (RL) for solving discounted Markov decision processes (MDPs) when the transition kernel is unknown.
By Donghwan Lee
arXiv:2511. 03836v2 Announce Type: replace Abstract: Deep Q-Networks (DQNs) estimate future returns by learning from transitions sampled from a replay buffer.
By Lipeng Zu, Hansong Zhou, Xiaonan Zhang
The paper introduces Q-Target Pretrained Transformers (QTPT), a method that replaces supervised behavior cloning with a Bellman-style Q‑target objective for in‑context reinforcement learning. QTPT retains the context‑conditioned Transformer architecture but learns to estimate action values using rewards and transitions from the context, rather than merely imitating offline actions. The authors provide theoretical analysis in stochastic linear bandits and finite‑horizon MDPs, demonstrating improved robustness to weak or suboptimal data, and empirically show gains over supervised pretraining on controlled RL benchmarks and extensions to D4RL Kitchen and AntMaze.
By Yichen Lin, Xuyuan Xiong, Xue Wang, Xiangfu Meng, Mike Mingcheng Wei, Tao Yao
The paper investigates hybrid quantum‑classical neural networks for learning routing heuristics, focusing on whether small quantum neural networks can replace parameter‑heavy modules in an attention‑based routing model without sacrificing solution quality. For the capacitated vehicle routing problem, replacing the encoder feed‑forward component with a quantum version reduces model parameters by 56.6% while maintaining performance close to the classical baseline on small and medium instances, though the gap widens for larger instances. The study also compares the hybrid approach to classical routing algorithms, finding that classical methods remain highly competitive and often superior on fixed Euclidean test sets, indicating no quantum advantage but highlighting encoder feed‑forward replacement as a viable compression strategy for neural combinatorial optimization.
By Marcus Rolf Peter Ritt, Alexsandro Santos da Rosa J\'unior, Marcos Vinicius Reballo, Cesar Augusto do Amaral, Fernando Augusto Caletti de Barros