Which Medical Questions Deserve Rationales? Perturbation-Sensitive Selection for Robust QA
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2609.24799v1 Announce Type: new Abstract: Post-training quantization (PTQ) enables efficient deployment of large language models, and PTQ methods are usually optimized and evaluated with generi...
The paper introduces Expected‑Severity‑Risk (ESR), a new objective for selecting questions in proactive medical dialogue that prioritizes reducing the expected severity of diagnostic errors rather than merely uncertainty. ESR uses population statistics to marginalize over possible answers and distills its rankings into a prefix‑only language policy, enabling deployment without teacher‑side risk computation. Experiments on DDxPlus show ESR cuts high‑severity diagnostic misses by 29.5% and boosts accuracy while adding only 0.14 extra questions per dialogue.
arXiv:2609.27987v1 Announce Type: new Abstract: Large language models (LLMs) have made substantial progress on medical question-answering, yet effective medical dialogue also requires learning to ask...
arXiv:2606. 05174v1 Announce Type: cross Abstract: Large Language Models (LLMs) have shown strong promise in healthcare applications.
The paper introduces TriageRA-CCF, a method for adaptively allocating low‑rank LoRA channels in medical large language models based on source‑side signals: answer confidence, clinical coverage, and a counterfactual close‑miss proxy. By supervising a budget router that selects among 2, 4, or 8 active ranks, the approach improves average accuracy over existing LoRA variants on Qwen3‑8B and Llama3.1‑8B, though gains vary across benchmarks. Ablation studies confirm that each signal contributes to better budget decisions, though their combined effect is not uniformly superior across all backbones.
arXiv:2606. 10279v1 Announce Type: new Abstract: Supervised fine-tuning with synthetic rationale data is widely assumed to improve language model performance on clinical prediction tasks by teaching models not just what to predict but why.