Evaluating Fine-Tuned and Base Language Models in Maternal and Vaccination Healthcare for African Settings
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2606. 08483v1 Announce Type: new Abstract: Background: Consumer-facing large language models are now a common source of health information, and they interpret and personalize responses rather than retrieve them.
MIRA is a bilingual benchmark that evaluates whether large language models (LLMs) provide consistent medical information across different user phrasings, languages, and health literacy levels. It contains 4,320 prompts derived from 60 medically reviewed low‑risk health questions and reveals that models tend to omit key information and offer fewer concrete next steps when responding to low health‑literacy signals, a phenomenon termed Differential Information Dilution (DID). A knowledge‑guided mitigation prompt can reduce this dilution for most models, notably improving Claude and Qwen.
arXiv:2608. 12138v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings.
arXiv:2604. 14892v3 Announce Type: replace-cross Abstract: Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators.
arXiv:2512. 01241v3 Announce Type: replace-cross Abstract: Large language models (LLMs) are routinely used by physicians and patients for medical advice, yet their clinical safety profiles remain poorly characterized.
arXiv:2601. 17642v2 Announce Type: replace Abstract: Safety alignment in Large Language Models is critical for healthcare; however, reliance on binary refusal boundaries often results in over-refusal of benign queries or unsafe compliance with harmful ones.