arXiv Machine Learning By Joshua P. Zitovsky, Yating Zou, Leslie Wilson, Michael R. Kosorok

Latent Utility Q-Learning for Preference-Adaptive Dynamic Treatment Regimes

Read the original on arXiv Machine Learning →

arXiv:2307. 12022v3 Announce Type: replace-cross Abstract: Optimizing individualized treatment sequences for patients who weigh multiple, competing outcomes differently poses a challenge for dynamic treatment regime (DTR) methods, which typically assume a single univariate outcome.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
1d ago

Patient-Centered Treatment Planning for Chronic Multimorbidity: A Hierarchical Reinforcement Learning Framework for Preference Modeling

The paper introduces FAHOC, a hierarchical reinforcement learning framework that models patient preferences by learning high‑level therapeutic options and factored intra‑option policies, while enforcing a cooperation‑aware action masking mechanism. It demonstrates that cooperative patients achieve better health outcomes and that the Q‑function approximation error is bounded. Evaluated on data from ~50,000 comorbid hypertension and type 2 diabetes patients, FAHOC improves quality‑adjusted life years by 0.669, correctly identifies cooperative patients 95.9% of the time, and never violates patient preferences in held‑out tests.

By Nafiseh Payani, Soham Das, G. Anthony Wilson, Anahita Khojandi
arXiv Machine Learning
Sep 10

Decision-Centered Abstractions via Orthogonal Estimation of Difference-of-Q Functions

The paper introduces state abstractions that preserve the difference of Q‑functions for offline reinforcement learning, aiming to exclude irrelevant dynamics from rich state data. It proposes a dynamic generalization of the R‑learner that uses orthogonal estimation and sparse learning to estimate the Q‑function contrast, achieving faster convergence and consistency under a margin condition. Experiments on simulated and simulator‑augmented real data show variance reductions and demonstrate that the necessary information for sequential decision‑making can be smaller than that required for full state prediction.

By Defu Cao, Angela Zhou