arXiv Machine Learning

SCOPE: Supply-Chain Operations through Coupled Policies for End-to-End Coordination

arXiv:2607. 28488v1 Announce Type: cross Abstract: Can supply-chain AI move beyond isolated decision modules toward unified operational planning?

arXiv AI
Jun 12

Multi-Agent Reinforcement Learning from Delayed Marketplace Feedback for Objective-Weight Adaptation in Three-Sided Dispatch

arXiv:2606. 13604v1 Announce Type: new Abstract: Dispatch in three-sided marketplaces provides a natural setting for reinforcement learning from world feedback: decisions are evaluated by delayed operational outcomes such as delivery speed, courier utilization, and merchant congestion.

By Haochen Wu, Yi Hou, Shiguang Xie
Hugging Face Trending Papers
Jun 11

Multi-Agent Reinforcement Learning from Delayed Marketplace Feedback for Objective-Weight Adaptation in Three-Sided Dispatch

Dispatch in three-sided marketplaces provides a natural setting for reinforcement learning from world feedback: decisions are evaluated by delayed operational outcomes such as delivery speed, courier utilization, and merchant congestion. We present a deployed reinforcement learning system at DoorDash that adapts dispatch objective weights in a large-scale food-delivery marketplace using delayed signals.

arXiv AI
Jun 18

Maturing Markov Decision Processes: Decision Making under Increasing Information and Shrinking Action Sets

arXiv:2606. 18820v1 Announce Type: cross Abstract: Sequential decision problems often exhibit an asymmetric evolution of information and decision flexibility: as a decision cycle unfolds, the agent receives richer information while feasible actions expire due to operational cutoffs, commitments, or resource constraints.

By Jiaxi Liu, Aiping Yang, Yuhang Yang, Shuqi Zhang, Zewei Dong, Jiangming Yang, Xuebin Chen
arXiv AI
Sep 4

Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations

The paper presents a graph‑constrained agentic framework that enables large language models to adapt retail supply‑chain decision modules to evolving requirements. It jointly selects intervention routes and admissible module changes, validating candidates against downstream KPIs. Experiments with 100 warehouse requirements and three LLMs show the framework improves end‑to‑end success from 72–76% to 79–83%.

By Lei Zheng, Liping Yang, Zihao Li, Guodong Lyu, Chaik Ming Koh, Chung-Piaw Teo
arXiv Machine Learning
Sep 3

OR-Transformer: Scaling Real-Time Decision-Making to 1,000 Items

OR-Transformer is a deep reinforcement learning framework designed for joint replenishment in supply chain operations with thousands of items. It uses a permutation‑equivariant Transformer architecture and pathwise‑gradient training to handle high‑dimensional observation and action spaces. In tests up to 1,024 items, it outperforms both learning‑based and rolling‑horizon MILP baselines and cuts online decision time by over four million times.

By Shuze Daniel Liu, David Simchi-Levi, Claire Chen, Chutong Gao, Shangtong Zhang
arXiv AI
Jul 21

A Deep Reinforcement Learning Algorithm for the Vehicle Routing Problem with Stochastic Demands and Outsourcing

arXiv:2607. 16875v1 Announce Type: cross Abstract: We introduce the vehicle routing problem with stochastic demands and outsourcing options (VRP-SDO), in which a logistics service provider partitions customer requests into customers outsourced to a common carrier and customers committed to its fixed fleet.

By Mohsen Dastpak, Fausto Errico, Ola Jabali