The paper proposes a mean‑field reinforcement learning framework that models rewards and transitions as functions of an unknown low‑dimensional aggregate statistic of a large agent population. By learning this low‑dimensional representation in an offline setting, the authors demonstrate a provable method for obtaining near‑optimal policies. Experiments on a one‑step routing game inspired by supply‑chain problems show that, with a fixed neural‑network size and optimization budget, the learned representation improves reward prediction and the quality of Nash equilibria compared to baselines that ignore population structure.
By Aditya Makkar, Benjamin Unger, Jeongyeol Kwon, Mathieu Lauri\`ere, Eugene Vinitsky, Yonathan Efroni
arXiv:2609.36174v1 Announce Type: new
Abstract: We consider fair resource allocation in sequential decision-making environments modeled as major-minor weakly coupled Markov decision processes (M2WCMD...
By Xiaohui Tu, Yossiri Adulyasak, Erick Delage
RideSkill is a hierarchical algorithm for generalized ride sharing that uses large language models (LLMs) to automatically design and train a skill repository, a combiner, and a repositioner. The combiner assigns vehicle-specific skills for adaptive dispatch across varying scenarios and objectives, while the repositioner moves idle vehicles to emerging regions to avoid conflicts. By training all components via an LLM-based evolutionary method, RideSkill eliminates the need for real-time LLM calls, enabling high-performance deployment in large-scale systems.
By Zijian Zhao, Sen Li, Xialiang Tong, Mingxuan Yuan
arXiv:2601. 18783v2 Announce Type: replace-cross Abstract: Balancing safety, efficiency, and operational costs in highway driving poses a challenging decision-making problem for heavy-duty vehicles.
By Deepthi Pathare, Leo Laine, Morteza Haghir Chehreghani
arXiv:2607. 18286v1 Announce Type: cross Abstract: Transit signal priority (TSP) requires balancing competing objectives: reducing bus delay while limiting adverse impacts on non-bus traffic and avoiding extreme waits for a subset of vehicles.
By Philip-Roman Adam, Stefanie Schmidtner
arXiv:2604. 25848v2 Announce Type: replace Abstract: We study city-scale control of electric-vehicle (EV) ride-hailing fleets where dispatch, repositioning, and charging decisions must respect charger and feeder limits under uncertain, spatially correlated demand and travel times.
By An Nguyen, Hoang Nguyen, Phuong Le, Hung Pham, Cuong Do, Laurent El Ghaoui
arXiv:2603. 06607v2 Announce Type: replace-cross Abstract: Radio resource allocation (RRA) is a critical function in cellular vehicle-to-everything (C-V2X) networks, where vehicles must share limited wireless resources to support safety-critical communications.
By Siyuan Wang, Lei Lei, Pranav Maheshwari, Sam Bellefeuille, Kan Zheng
arXiv:2607. 16875v1 Announce Type: cross Abstract: We introduce the vehicle routing problem with stochastic demands and outsourcing options (VRP-SDO), in which a logistics service provider partitions customer requests into customers outsourced to a common carrier and customers committed to its fixed fleet.
By Mohsen Dastpak, Fausto Errico, Ola Jabali
arXiv:2606. 18111v1 Announce Type: cross Abstract: Fairness is an important aspect of decision-making in multi-objective reinforcement learning (MORL), where policies must ensure both optimality and equity across multiple, potentially conflicting objectives.
By Umer Siddique, Peilang Li, Yongcan Cao
CoRe-MARL is a cooperative multi-agent reinforcement learning framework designed for decentralized relief distribution networks. It models each local center as an agent in a Dec-POMDP, using a recurrent network to learn redistribution policies that reduce service gaps and improve the worst-served region. Experiments show that recurrent MAPPO outperforms independent PPO and heuristic baselines, maintaining competitive network-wide service while adapting to evolving supply and demand dynamics.
By Naimur Rahman Chowdhury, Shatabdi Sen Prapti, Md. Salehin Seyam, Limon Bin Hossain
FedPGT introduces a progressive gradient transmission scheme for vehicular federated learning over time‑varying channels, where vehicles send high‑magnitude gradient entries according to instantaneous channel conditions. The authors derive a convergence bound showing diminishing returns governed by a power‑law decay, and formulate a stochastic optimization problem that is solved via a Lyapunov drift‑plus‑penalty approach with per‑slot surrogate variables. A low‑complexity resource allocation algorithm is proposed, and experiments on CIFAR‑10 and Argoverse demonstrate a 3.65% accuracy gain and a 12.66% reduction in displacement error compared to state‑of‑the‑art baselines.
By Jintao Yan, Tan Chen, Yuxuan Sun, Sheng Zhou, Zhisheng Niu
Emergency management assistance programs, such as relief distribution, are essential for delivering necessary supplies to affected communities. However, these programs operate in a decentralized netwo...