arXiv Machine Learning By Xianhua Peng, Wu Guo

Deep Learning for Dynamic Programming with Recursive Utility

Read the original on arXiv Machine Learning →

arXiv:2607. 04278v1 Announce Type: cross Abstract: We propose the first deep learning algorithm, the Certainty Equivalent Learning (CEL) algorithm, for solving high-dimensional discrete-time dynamic programming problems with recursive utility.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
5d ago

Learning Chance-Constrained MDPs with Bellman Distributional Certificates

The paper introduces a new approach to learning chance-constrained Markov decision processes (CCMDPs) using a Bellman distributional certificate. It provides both model-based and model-free algorithms with theoretical guarantees, including matching upper and lower bounds for tabular discounted CCMDPs with bounded successor support. Numerical experiments on synthetic CCMDPs and an IEEE 14-bus energy storage benchmark demonstrate the safety and effectiveness of the proposed methods.

By Chenbei Lu, Hongyu Yi
arXiv AI
Sep 17

Deep Learning for Sequential Decision Making under Uncertainty: Foundations, Frameworks, and Frontiers

The tutorial titled "Deep Learning for Sequential Decision Making under Uncertainty: Foundations, Frameworks, and Frontiers" explores how modern deep learning techniques—such as neural networks, transformers, large language models, and deep reinforcement learning—can be integrated with operations research and management science to address complex, uncertain, and dynamic decision problems. It argues that deep learning should complement, not replace, optimization, offering adaptability and scalable approximation while OR/MS provides rigorous constraint and uncertainty modeling. The tutorial organizes the field around predict‑then‑optimize, decision‑aware learning, constraint‑aware decision generation, and deep reinforcement learning, and highlights applications across supply chains, healthcare, energy, and autonomous systems.

By I. Esra Buyuktahtakin
arXiv Machine Learning
Jul 17

A Noise-Robust Elicit-to-Optimize Framework for Distortion Riskmetrics via Inverse Reinforcement Learning

arXiv:2607. 14373v1 Announce Type: new Abstract: We propose a noise-robust elicit-to-optimize framework that integrates inverse reinforcement learning (IRL) and reinforcement learning (RL) for eliciting agents' risk preferences and optimizing policies under a broad class of risk objectives characterized by distortion riskmetrics.

By Yang Liu, Yuhao Liu, Yunran Wei