arXiv Machine Learning By Hibat Errahmen Djecta, Sergey Alyaev, Kristian Fossum, Reidar B. Bratvold, Ressi Bonti Muhammad, Apoorv Srivastava

Decision-Driven Geosteering Under Uncertainty: A Unified Framework for Sequential Decision Optimization

Read the original on arXiv Machine Learning →

arXiv:2606. 17331v1 Announce Type: new Abstract: Geosteering requires navigating a well trajectory through an unknown geological configuration, while sequentially updating decisions based on indirect measurements acquired during drilling.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 17

Deep Learning for Sequential Decision Making under Uncertainty: Foundations, Frameworks, and Frontiers

The tutorial titled "Deep Learning for Sequential Decision Making under Uncertainty: Foundations, Frameworks, and Frontiers" explores how modern deep learning techniques—such as neural networks, transformers, large language models, and deep reinforcement learning—can be integrated with operations research and management science to address complex, uncertain, and dynamic decision problems. It argues that deep learning should complement, not replace, optimization, offering adaptability and scalable approximation while OR/MS provides rigorous constraint and uncertainty modeling. The tutorial organizes the field around predict‑then‑optimize, decision‑aware learning, constraint‑aware decision generation, and deep reinforcement learning, and highlights applications across supply chains, healthcare, energy, and autonomous systems.

By I. Esra Buyuktahtakin
arXiv Machine Learning
Aug 24

Smart Exploration in Reinforcement Learning using Bounded Uncertainty Models

The paper introduces BUMEX, a reinforcement learning exploration strategy that leverages a set of prior models containing the true transition kernel and reward function. By optimizing over this model set, the method derives upper and lower bounds on the Q‑function to guide exploration, providing theoretical guarantees of convergence to the optimal policy. When the model set follows a bounded‑parameter MDP structure, the optimization becomes convex, enabling finite‑time convergence under mild assumptions and demonstrating accelerated learning in simulations.

By J. S. van Hulst, W. P. M. H. Heemels, D. J. Antunes