Dispatch in three-sided marketplaces provides a natural setting for reinforcement learning from world feedback: decisions are evaluated by delayed operational outcomes such as delivery speed, courier utilization, and merchant congestion. We present a deployed reinforcement learning system at DoorDash that adapts dispatch objective weights in a large-scale food-delivery marketplace using delayed signals.
arXiv:2608. 10897v1 Announce Type: new Abstract: Instant delivery platforms have become a critical component of urban logistics, increasingly relying on crowdsourced couriers to fulfill highly dynamic orders.
By Fengming Yao, Man Luo
The paper introduces proximal residual value functions for two‑timescale decision systems, where a planning layer supplies a continuation‑value function to a real‑time optimizer that allocates resources, with inventory placement as a motivating example. The authors propose an end‑to‑end reinforcement learning method that learns a convex residual added to a strictly convex potential, enabling well‑posed optimization and end‑to‑end differentiation while maintaining an explicit convex objective for real‑time execution. They also provide necessary and sufficient conditions for smooth value functions to produce decisions consistent across planning and execution timescales, and demonstrate a 5.0% reduction in routing and transfer cost in an offline simulation using data from a large e‑commerce retailer.
By Harrison Waldon, Carson Eisenach, Akhil Bagaria, Daniel Russo, Dominique Perrault-Joncas, Alisha Zachariah, Dean Foster
arXiv:2607. 22356v1 Announce Type: new Abstract: In recent years, the growing complexity of last-mile pickup operations has increased the need for fast and accurate decision-making on logistics platforms.
By Yida Xu, Zhaofang Mao, Yuheng Miao, Jiaxin Zhang, Yiting Sun
arXiv:2608.28878v1 Announce Type: cross
Abstract: This paper develops a hybrid offline-online multi-agent reinforcement learning framework based on decision transformers. The policy is first pretrain...
By Yiming Zhang, Kun Yang, Cong Shen, Dongning Guo
arXiv:2607. 16875v1 Announce Type: cross Abstract: We introduce the vehicle routing problem with stochastic demands and outsourcing options (VRP-SDO), in which a logistics service provider partitions customer requests into customers outsourced to a common carrier and customers committed to its fixed fleet.
By Mohsen Dastpak, Fausto Errico, Ola Jabali
arXiv:2607. 28488v1 Announce Type: cross Abstract: Can supply-chain AI move beyond isolated decision modules toward unified operational planning?
By Yunhao Liang, Xianqi Cao, Pujun Zhang, Yuan Qu, Yongzhi Qi, Ningxuan Kang, Max Z. J. Shen
arXiv:2606. 06201v1 Announce Type: new Abstract: Pharmaceutical supply chains (PSCs) struggle with inventory management (IM) due to unpredictable demand patterns and variable lead times associated with restocking.
By Amandeep Kaur, Gyan Prakash
arXiv:2501. 12942v2 Announce Type: replace Abstract: Effective multi-user delay-constrained scheduling is crucial in various real-world applications, including embodied AI, instant messaging, live streaming, and data center management, where efficient resource allocation is required among users with diverse delay sensitivities.
By Zhuoran Li, Ruishuo Chen, Hai Zhong, Longbo Huang
arXiv:2601. 02754v3 Announce Type: replace-cross Abstract: With the rapid development of e-commerce, auto-bidding has become a key asset in optimizing advertising performance under diverse advertiser environments.
By Mingming Zhang, Na Li, Zhuang Feiqing, Hongyang Zheng, Jiangbing Zhou, Wang Wuyin, Sheng-jie Sun, XiaoWei Chen, Junxiong Zhu, Lixin Zou, Chenliang Li
arXiv:2606. 23978v1 Announce Type: cross Abstract: We present an offline reinforcement learning (RL) framework for optimizing SLAM throughput control in a warehouse fulfillment environment.
By Tina Dongxu Li, Mouhacine Benosman, Rajat Kumar, Kevin Tan, Ken Meszaros, Trevor Dardik
The paper proposes a deployment‑focused framework for deadline‑constrained network control, introducing the Effective Congestion (EC) metric family and Uniform Path Grouping (UPG) heuristic to better capture traffic urgency and balance load. It integrates these with a Multi‑Agent Deep Reinforcement Learning architecture (MADRL EC (p*)) that combines a distributed scheduler and a centralized RL router. A unified training objective merges live‑reward, pre‑collected‑reward, and policy‑imitation terms, leading to the Model‑Guided Annealed Reinforcement Learning (MGA‑RL) protocol built on DDPG, which generalizes offline‑to‑online learning for demonstration‑driven training.
By Vincenzo Norman Vitale, Mohammad Solki, Antonia Maria Tulino, Andreas F. Molisch, Jaime Llorca