The paper introduces an evidence ladder for evaluating reinforcement learning (RL) in healthcare, outlining stages from problem formulation to lifecycle monitoring. It argues that success in historical data does not guarantee real‑world improvement and highlights assumptions and failure modes at each rung. The authors propose reporting practices to support cumulative evaluation and emphasize that RL should be tested as an intervention within a dynamic sociotechnical system.
By Yunfan Zhao
arXiv:2605. 19208v2 Announce Type: replace-cross Abstract: Physical activity (PA) plays an important role in maintaining and improving health.
By Gefei Lin, Rui Miao, Jennifer Sacheck, Xiaoke Zhang
arXiv:2508. 03875v2 Announce Type: replace Abstract: Many sequential decision problems offer qualitatively different ways of influencing the environment: some interventions act immediately, whereas others induce persistent effects that continue to shape future states long after the decision that initiated them.
By David Mguni, Wanrong Yang, Jing Dong, Ziquan Liu, Muhammad Salman Haleem, Baoxiang Wang, Dominik Wojtczak
arXiv:2608.22615v1 Announce Type: new
Abstract: Large Language Model (LLM)-based counseling agents can generate fluent and supportive responses, but they often lack the structured, goal-directed prog...
By Qi Zhang, Heajun An, Prakriti Dumaru, Sang Won Lee, Lifu Huang, Pamela J. Wisniewski, Jin-Hee Cho
arXiv:2601. 15353v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) has achieved remarkable success in real-world decision-making across diverse domains, including gaming, robotics, online advertising, public health, and natural language processing.
By Asim H. Gazi, Yongyi Guo, Daiqi Gao, Ziping Xu, Kelly W. Zhang, Susan A. Murphy
The paper introduces POROS, a framework that uses peer‑grounded counterfactual explanations to generate incremental behavioral steps for chronic disease management. POROS builds a directed acyclic graph of patient states where each edge represents a behavior change that peers have successfully achieved and that improves health outcomes. In two diabetes cohorts, POROS reduces the required improvement per step from over 25–30 percentage points to about 5–6, while most multi‑hop paths involve cross‑patient comparisons.
By Saman Khamesian, Hassan Ghasemzadeh