arXiv:2608. 02617v1 Announce Type: cross Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform collecting blinded pairwise preferences alongside multi-criterion rubric ratings.
By Fay Elhassan, David Sasu, Alexandra Kulinkina, Lars Henning Klein, Mary-Anne Hartley
The study evaluates how quantization affects accuracy and safety of five 7‑8B language models on clinical benchmarks. INT8 GPTQ shows minimal degradation (≤1.9%) across tasks, while INT4 causes substantial, model‑dependent drops, especially in high‑risk scenarios and safety metrics. Recovery methods such as clinical calibration substitution and QLoRA fine‑tuning yield mixed results, underscoring the need for task‑specific validation.
By Leonard Twagirayezu, Prasenjit Mitra
The paper introduces a retrieval‑augmented multi‑agent framework that automatically generates instance‑specific evaluation rubrics for medical language models. By retrieving authoritative medical evidence, decomposing it into atomic facts, and combining these with user interaction constraints, the system produces fine‑grained criteria that outperform GPT‑4o on HealthBench and LLMEval‑Med. The generated rubrics also guide response refinement, improving medical LLM output quality by 9.2%.
By Yinzhu Chen, Abdine Maiga, Hossein A. Rahmani, Emine Yilmaz
arXiv:2608. 04772v1 Announce Type: cross Abstract: Scaling supervision for multi-turn medical agents is difficult because expert dialogue annotation is costly and clinical conversations are privacy-restricted.
By Chenyu Wang, Yi Liu, Baoqing Li, Min Tu, Diping Song
arXiv:2607. 18086v1 Announce Type: new Abstract: Background: LLM judges increasingly score whether clinical language models give overconfident answers under incomplete evidence, yet whether a measured "safety gain" reflects real behavior change or the judge's calibration is unresolved.
By Koyar Afrasyab
arXiv:2607. 20453v1 Announce Type: cross Abstract: Large language models show promise for clinical prediction, but zero-shot performance on specialized tasks is limited by incomplete domain knowledge, especially for smaller locally deployable models.
By Jessica Sena, Shesadree Priyadarshani, Miguel Contreras, Bharat Gandhi, Scott Siegel, Subhash Nerella, Parisa Rashidi
arXiv:2608. 03854v1 Announce Type: new Abstract: When decoder language models are used as classifiers, predicted class probabilities depend on implementation choices, including the prompt template, verbalizer (label-to-token mapping), and scoring rule, that are rarely treated as experimental variables.
By Anton Rasmussen, Hong Qin
The paper investigates how to incorporate biomedical knowledge graphs (KGs) into large language models (LLMs) for clinical diagnosis. It evaluates five KG task formulations, three training paradigms, two KGs, and three base LLMs, finding that all paradigms outperform a non‑finetuned baseline but differ in knowledge transfer behavior. Introducing Gradient Intervention Density (GID) and Gradient Distortion (GD) metrics, the study identifies a ‘surgical alignment’ regime—sparse, localized updates achieved by KG‑judgment training with KL regularization—that improves reasoning quality even when in‑domain accuracy is lower than task‑specific supervised fine‑tuning.
By Saksham Khatwani, He Cheng, Majid Afshar, Dmitriy Dligach, Yanjun Gao
arXiv:2607. 24371v1 Announce Type: cross Abstract: Healthcare interoperability requires AI systems to produce structured outputs conforming to standardized schemas including ICD-10 for diagnostic coding, CPT for procedure billing, and HL7 FHIR for data exchange.
By Jianru Shen
arXiv:2606. 28332v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for medical and health-related questions, yet their safety in high-risk medical scenarios remains poorly understood.
By Yige Li, Jun Sun, Wei Zhao, Zhe Li, Yutao Wu, Hanxun Huang, Xiang Zheng, Xingjun Ma
arXiv:2601. 17642v2 Announce Type: replace Abstract: Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-refusal of benign queries or unsafe compliance with harmful ones.
By Zhihao Zhang, Liting Huang, Guanghao Wu, Preslav Nakov, Heng Ji, Usman Naseem
arXiv:2606. 30887v1 Announce Type: cross Abstract: Large language models show promise for mental health support, yet therapeutic quality improves only when evaluation functions as an actionable control signal rather than a passive metric.
By Mizanur Rahman, Abeer Badawi, Elahe Rahimi, Laleh Seyyed-Kalantari, Frank Rudzicz, Enamul Hoque, Elham Dolatabadi