Recent offline reinforcement learning (RL) studies report policies that outperform physician decisions on clinical outcomes. We conduct a systematic, partially crossed evaluation of five offline RL al...
arXiv:2608. 03606v1 Announce Type: new Abstract: Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneous evidence.
By William Bolton, Philip Torr
arXiv:2606. 05932v1 Announce Type: cross Abstract: Reinforcement learning from verifiable rewards (RLVR) improves reasoning even when the reward signal is spurious -- assigning credit to the group-plurality answer rather than a ground-truth verifier.
By Yuze Gao
The study develops a 1‑D ResNet classifier that uses a fixed 17‑channel hemodynamic representation derived from photoplethysmography (PPG) to predict in‑hospital stroke risk states up to six hours before clinical recognition. Using data from MIMIC‑III and MC‑MED, the model achieved F1‑scores ranging from 0.7956 to 0.9888 across 4‑, 5‑, and 6‑hour horizons, outperforming four non‑waveform clinical and structured‑EHR comparators in all cohort‑horizon settings. Retrospective analysis showed that the PPG model’s false‑positive rates on high‑risk non‑stroke controls could be reduced through persistence aggregation, though the study does not establish a calibrated bedside alarm or a clinically validated prediction lead time.
By Jiaming Liu, Cheng Ding, Jian Wu, Hongxia Xu, Daoqiang Zhang
arXiv:2609. 31450v1 Announce Type: cross Abstract: Medical vision-language model (VLM) post-training is commonly evaluated through answer accuracy.
By Wang Jingxin
arXiv:2608. 05203v1 Announce Type: new Abstract: Machine learning models achieve strong predictive accuracy for 90-day outcome prediction in acute ischaemic stroke, yet clinical adoption is limited by the misalignment of model explanations with clinicians' reasoning.
By Esra Zihni, Katryna Cisek, Hamzah Ziadeh, Hendrik Knoche, Robert Mikulik, John D. Kelleher