In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate this issue, behavior policy search has been propo...
arXiv:2602. 17375v3 Announce Type: replace Abstract: We formulate episodic Markov decision process (MDP) planning as Bayesian inference over policies.
By David Tolpin
arXiv:2512. 01362v2 Announce Type: replace Abstract: Neural prediction offers a promising approach to forecasting the individual variability of neurocognitive functions and disorders and providing prognostic indicators for personalized invention.
By Yanlin Wang, Nancy M Young, Patrick C M Wong
arXiv:2603. 24705v3 Announce Type: replace-cross Abstract: Discrete choice models are fundamental tools in management science, economics, and marketing for understanding and predicting decision-making.
By Easton Huch, Michael Keane
arXiv:2608. 04669v1 Announce Type: new Abstract: Many social services assign scarce resources, such as housing assistance or hospital interventions, to people who arrive one at a time: each arrival must receive a decision immediately, and the long-run usage of every resource must stay within its capacity.
By Mohammadsaeed Haghi, Mahdi Salmani, Nima Kelidari
arXiv:2606. 18111v1 Announce Type: cross Abstract: Fairness is an important aspect of decision-making in multi-objective reinforcement learning (MORL), where policies must ensure both optimality and equity across multiple, potentially conflicting objectives.
By Umer Siddique, Peilang Li, Yongcan Cao
arXiv:2607. 02206v1 Announce Type: cross Abstract: Predictions are increasingly used to guide high-stakes decisions, from treatment selection to policy making.
By Yurui Zheng, Ying Jin
arXiv:2608. 08743v1 Announce Type: cross Abstract: Reinforcement learning (RL) seeks to optimize sequential decisions to maximize population-level benefits over time.
By Jianhan Zhang, Jitao Wang, John D. Piette, Donglin Zeng, Chengchun Shi, Zhenke Wu
arXiv:2601. 19612v3 Announce Type: replace-cross Abstract: Safe exploration is a key requirement for reinforcement learning (RL) agents to learn and adapt online, beyond controlled (e.
By Manuel Wendl, Yarden As, Manish Prajapat, Anton Pollak, Stelian Coros, Andreas Krause
arXiv:2608.24146v1 Announce Type: new
Abstract: In reinforcement learning policy evaluation, classic on-policy methods often suffer from high variance when estimating policy performance. To mitigate...
By Claire Chen, Shuze Daniel Liu, Licheng Luo, Rohan Chandra, Nan Jiang, Shangtong Zhang
arXiv:2606. 18438v1 Announce Type: cross Abstract: In this paper, we study a sequential workforce management problem in a contingent labor setting with uncertainty in both worker production and labor supply.
By Chris Lee, Xiuli Chao, Izak Duenyas
arXiv:2608. 02893v1 Announce Type: cross Abstract: Counterfactual inference approaches for sequential decision-making typically assume deterministic causal models, where all randomness stems from latent variables.
By Jessica Lally, Milad Kazemi, Nicola Paoletti, David Watson, Sander Beckers