arXiv:2608.30442v1 Announce Type: new
Abstract: Recent offline reinforcement learning (RL) studies report policies that outperform physician decisions on clinical outcomes. We conduct a systematic, p...
By Kihun Rhee
arXiv:2608. 03606v1 Announce Type: new Abstract: Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneous evidence.
By William Bolton, Philip Torr
arXiv:2606. 05932v1 Announce Type: cross Abstract: Reinforcement learning from verifiable rewards (RLVR) improves reasoning even when the reward signal is spurious -- assigning credit to the group-plurality answer rather than a ground-truth verifier.
By Yuze Gao
arXiv:2609. 31450v1 Announce Type: cross Abstract: Medical vision-language model (VLM) post-training is commonly evaluated through answer accuracy.
By Wang Jingxin
arXiv:2608. 05203v1 Announce Type: new Abstract: Machine learning models achieve strong predictive accuracy for 90-day outcome prediction in acute ischaemic stroke, yet clinical adoption is limited by the misalignment of model explanations with clinicians' reasoning.
By Esra Zihni, Katryna Cisek, Hamzah Ziadeh, Hendrik Knoche, Robert Mikulik, John D. Kelleher
arXiv:2606. 05263v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards improves reasoning and tool use, yet long-horizon language agents still learn unsupported evidence chains, belief drift, and shortcut actions that satisfy terminal checks.
By Renwei Meng
arXiv:2608. 16482v1 Announce Type: new Abstract: The dosing of intravenous fluids and vasopressors in sepsis is a sequential decision made under uncertainty and guided largely by clinical judgment, which makes it a natural target for reinforcement learning from historical care.
By Marc P\'erez-Roig, David Fern\'andez-Narro, Carlos S\'aez
The dosing of intravenous fluids and vasopressors in sepsis is a sequential decision made under uncertainty and guided largely by clinical judgment, which makes it a natural target for reinforcement learning from historical care. Because a learned policy cannot be trialed on patients, its value must be estimated off-policy, and such estimates can be fragile and optimistic.
The paper introduces an evidence ladder for evaluating reinforcement learning (RL) in healthcare, outlining stages from problem formulation to lifecycle monitoring. It argues that success in historical data does not guarantee real‑world improvement and highlights assumptions and failure modes at each rung. The authors propose reporting practices to support cumulative evaluation and emphasize that RL should be tested as an intervention within a dynamic sociotechnical system.
By Yunfan Zhao
arXiv:2510. 19893v2 Announce Type: replace Abstract: Medical AI systems demonstrated impressive diagnostic performance, yet they routinely show uneven accuracy across demographic groups, disadvantaging underrepresented populations.
By Shiqi Dai, Wei Dai, Jiaee Cheong, Paul Pu Liang
arXiv:2606. 20206v1 Announce Type: cross Abstract: In offline Reinforcement Learning, immediate rewards in logged batch data are often unobserved due to sparse or irregular record-keeping, or censored beyond certain reward values.
By Ziheng Wei, Annie Qu, Rui Miao
The study develops a 1‑D ResNet classifier that uses a fixed 17‑channel hemodynamic representation derived from photoplethysmography (PPG) to predict in‑hospital stroke risk states up to six hours before clinical recognition. Using data from MIMIC‑III and MC‑MED, the model achieved F1‑scores ranging from 0.7956 to 0.9888 across 4‑, 5‑, and 6‑hour horizons, outperforming four non‑waveform clinical and structured‑EHR comparators in all cohort‑horizon settings. Retrospective analysis showed that the PPG model’s false‑positive rates on high‑risk non‑stroke controls could be reduced through persistence aggregation, though the study does not establish a calibrated bedside alarm or a clinically validated prediction lead time.
By Jiaming Liu, Cheng Ding, Jian Wu, Hongxia Xu, Daoqiang Zhang