The paper introduces a retrieval‑augmented multi‑agent framework that automatically generates instance‑specific evaluation rubrics for medical language models. By retrieving authoritative medical evidence, decomposing it into atomic facts, and combining these with user interaction constraints, the system produces fine‑grained criteria that outperform GPT‑4o on HealthBench and LLMEval‑Med. The generated rubrics also guide response refinement, improving medical LLM output quality by 9.2%.
By Yinzhu Chen, Abdine Maiga, Hossein A. Rahmani, Emine Yilmaz
arXiv:2606. 05970v1 Announce Type: cross Abstract: Large language models are increasingly used for structured extraction from clinical free-text notes, but the sensitivity of their output to upstream configuration choices is less understood than their accuracy on fixed benchmarks.
By Martin Murin
arXiv:2606. 12569v1 Announce Type: cross Abstract: We present EDEN (Emergency Department Electronic Notes), a new and unique large-scale corpus of clinical notes produced in Emergency Departments of Italian hospitals.
By Tiziano Labruna, Guido Bertolini, Pietro Ferrazzi, Bernardo Magnini
arXiv:2605.11533v4 Announce Type: replace
Abstract: Routine clinical check-up reports combine laboratory measurements, physiological assessments, imaging findings and visually structured information,...
By Sike Xiang, Shuang Chen, Kevin Qinghong Lin, Jialin Yu, Yijia Sun, Philip Torr, Amir Atapour-Abarghouei
arXiv:2606. 12569v3 Announce Type: replace-cross Abstract: We present eCREAM-MedCorpus, a new and unique large-scale dataset of clinical notes produced in Emergency Departments of Italian hospitals.
By Tiziano Labruna, Guido Bertolini, Pietro Ferrazzi, Bernardo Magnini
arXiv:2606. 12569v2 Announce Type: replace-cross Abstract: We present eCream-MedCorpus, a new and unique large-scale dataset of clinical notes produced in Emergency Departments of Italian hospitals.
By Tiziano Labruna, Guido Bertolini, Pietro Ferrazzi, Bernardo Magnini