arXiv AI By Tina Dongxu Li, Mouhacine Benosman, Rajat Kumar, Kevin Tan, Ken Meszaros, Trevor Dardik

Offline Reinforcement Learning for Warehouse SLAM Throughput Control

Read the original on arXiv AI →

arXiv:2606. 23978v1 Announce Type: cross Abstract: We present an offline reinforcement learning (RL) framework for optimizing SLAM throughput control in a warehouse fulfillment environment.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 11

HiRAD: A Flexible Large-Scale AGV Routing System

HiRAD is a hierarchical reinforcement learning framework designed for continuous-space routing of large-scale AGV fleets, offering real-time guarantees. It introduces a step-level spatiotemporal representation, separates heading selection from velocity control to shrink the action space, and employs an asynchronous event-driven decision pipeline that reduces inference complexity from O(n²) to O(n) and cuts per-step latency by up to 71%. Experiments on random graphs and two warehouse maps show that HiRAD decreases makespan by 45% to 63% and shortens overall runtime.

By Yunjie Huang, Ruizhong Wu, Mengxuan Zhang, Frodo Kin Sun Chan, Yan Nei Law, Lei Li
arXiv AI
Aug 7

Search-Aided Joint Agent-Environment Reinforcement Learning for Robust Lifelong Multi-Agent Path Finding with Rotations

arXiv:2608. 05588v1 Announce Type: cross Abstract: Lifelong Multi-Agent Path Finding (LMAPF) requires repeatedly planning collision-free paths for agents that continuously receive new goals upon reaching their current ones.

By He Jiang, Jingtian Yan, Yulun Zhang, Yimin Tang, Tanishq Duhan, Rishi Veerapaneni, Guillaume Sartoretti, Jiaoyang Li
arXiv Machine Learning
Sep 4

Multi-step Proximal Policy Improvement in Offline Reinforcement Learning

The paper introduces Multi-step Proximal Policy Improvement (MPI), a method that refines offline reinforcement learning policies through sequential re-centered proximal steps. By viewing policies as a probability manifold, MPI interprets a wide range of offline actor objectives as a single proximal policy improvement step and extends this to multiple steps for controlled policy improvement beyond the behavior distribution. Experiments on D4RL benchmarks demonstrate that a few MPI refinements enhance strong offline baselines such as TD3+BC, ReBRAC, and IQL, while diagnostics clarify the benefits of re-centered refinement over fixed-objective scheduling and highlight critic error limitations.

By Soohyun Choi, Seonvin Cho, Songnam Hong