The paper introduces FAHOC, a hierarchical reinforcement learning framework that models patient preferences by learning high‑level therapeutic options and factored intra‑option policies, while enforcing a cooperation‑aware action masking mechanism. It demonstrates that cooperative patients achieve better health outcomes and that the Q‑function approximation error is bounded. Evaluated on data from ~50,000 comorbid hypertension and type 2 diabetes patients, FAHOC improves quality‑adjusted life years by 0.669, correctly identifies cooperative patients 95.9% of the time, and never violates patient preferences in held‑out tests.
By Nafiseh Payani, Soham Das, G. Anthony Wilson, Anahita Khojandi
arXiv:2607. 28916v1 Announce Type: cross Abstract: Multistep credit assignment is critical for sample-efficient reinforcement learning, yet managing off-policy bias in Q-learning remains a fundamental challenge.
By Brett Daley
arXiv:2606. 18111v1 Announce Type: cross Abstract: Fairness is an important aspect of decision-making in multi-objective reinforcement learning (MORL), where policies must ensure both optimality and equity across multiple, potentially conflicting objectives.
By Umer Siddique, Peilang Li, Yongcan Cao
arXiv:2605. 31014v2 Announce Type: replace Abstract: Multi-omics data provide complementary molecular characterizations of disease phenotypes and play an important role in disease diagnosis and subtype classification in precision medicine.
By Nan Mu, Yangfan Xiao, Ling Wang, Xiaoning Li, Yue Kang, Chen Zhao
The paper introduces state abstractions that preserve the difference of Q‑functions for offline reinforcement learning, aiming to exclude irrelevant dynamics from rich state data. It proposes a dynamic generalization of the R‑learner that uses orthogonal estimation and sparse learning to estimate the Q‑function contrast, achieving faster convergence and consistency under a margin condition. Experiments on simulated and simulator‑augmented real data show variance reductions and demonstrate that the necessary information for sequential decision‑making can be smaller than that required for full state prediction.
By Defu Cao, Angela Zhou
arXiv:2606. 23603v2 Announce Type: replace Abstract: Unhealthy dietary behavior continues to be a persistent public health issue in the United States, exacerbated by recommendation systems that prioritize user preference without considering nutritional health.
By Aarya Vasantlal, Joshua Zolla, Chuxu Zhang
arXiv:2607. 16916v1 Announce Type: new Abstract: Bladder cancer treatment requires personalized and adaptive decision-making, particularly for recurrent disease, where treatment effectiveness changes across successive clinical episodes.
By Divyansh Chawla, Anshu Garg, Isshaan Singh
arXiv:2606. 01051v1 Announce Type: new Abstract: Dynamic medical treatment requires deciding treatment intensity and intervention timing, while patient states evolve continuously and adverse events may occur between clinical interactions.
By Xun Shen, Yuepeng Wang, Akifumi Wachi, Yongqi Zhou, Richard Weiss, Yoshihiko Fujisawa, Ken Kawano, Mehrshad Sadria, Ying Chen, Xin Liu, Sebastien Gros, Xiao Hu, Kyoung-Sook Kim, Mengmou Li, Katsuki Fujisawa, Kenji Wakabayashi
arXiv:2506. 13702v4 Announce Type: replace-cross Abstract: Single-trajectory preference optimization methods learn from datasets of ((prompt, response, reward)) tuples, offering a practical alternative to pairwise preference learning by directly leveraging scalar feedback.
By Bilal Faye, Hanane Azzag, Mustapha Lebbah
arXiv:2508. 03875v2 Announce Type: replace Abstract: Many sequential decision problems offer qualitatively different ways of influencing the environment: some interventions act immediately, whereas others induce persistent effects that continue to shape future states long after the decision that initiated them.
By David Mguni, Wanrong Yang, Jing Dong, Ziquan Liu, Muhammad Salman Haleem, Baoxiang Wang, Dominik Wojtczak
arXiv:2609.36390v1 Announce Type: cross
Abstract: Offline reinforcement learning seeks optimal decision rules from previously collected data. In some applications, a decision can be an entire functio...
By Gefei Lin, Rui Miao, Xiaoke Zhang
Pre-training followed by fine-tuning has become the dominant recipe for learning performant policies, and in value-based reinforcement learning (RL) this raises a natural question: given a pretrained policy, should the Q-function be pretrained on offline data too? Conventional wisdom suggests it should, but recent results show that online RL with a randomly-initialized Q-function can result in highly performant and reliable policies without needing to pretrain the Q-function.