Hugging Face Trending Papers

MS-Exam-Gen: Source-Grounded Benchmark Construction for Evaluating LLMs on Textual Multiple Sclerosis MRI Knowledge

Read the original on Hugging Face Trending Papers →

MS-Exam-Gen is a reproducible framework that builds a source‑grounded multiple‑choice question benchmark for evaluating large language models on knowledge about multiple sclerosis MRI. The pipeline uses expert‑indexed sources, topic induction, evidence‑grounded MCQ generation, automated quality audits, and consistency checks to produce a 3,058‑item benchmark covering 16 topics and 53 subtopics. Evaluation of 12 LLM endpoints on this benchmark revealed a wide accuracy range (89.7% to 46.9%) and identified items frequently missed by models, while audits showed reduced answer cues and position‑sensitivity in scoring.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Jun 9

Automatic Extraction of Structured Information from Brain MRI Reports Using an Open-Weight Large Language Model

arXiv:2606. 07721v1 Announce Type: new Abstract: Objectives: Automatic data extraction from free-text radiology reports enables large-scale research, but few studies assessed the performance of large language models (LLMs) on Dutch neuroradiology reports.

By Kaouther Mouheb, Amos Pomp, Antoine Manenti, Romy de Haan, Farog Faghir, Joy Martens, Harro Seelaar, Francesco Mattace-Raso, Meike W. Vernooij, Frank J. Wolters, Stefan Klein, Esther E. Bron
arXiv AI
Aug 24

An integrated diffusion-weighted imaging processing and interpretation platform for MR-guided radiotherapy

An integrated, web‑based platform has been developed to process diffusion‑weighted imaging (DWI) from MR‑guided linear accelerators and provide structured, literature‑grounded clinical interpretations. The system combines a deep‑learning pipeline for distortion correction, denoising, and IVIM/ADC fitting with a retrieval‑augmented generation (RAG) agent that references a curated knowledge base and traces each statement to its source. Independent expert ratings of nine glioblastoma cases showed high scores for clinical reasoning, citation quality, and overall utility, with a mean rating of 4.65 out of 5.

By Yunxiang Li, Yan Dai, Yen-Peng Liao, Jie Deng, Jill B De Vis, You Zhang
arXiv AI
Jun 15

Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?

arXiv:2604. 14892v3 Announce Type: replace-cross Abstract: Evaluating medical AI systems using expert clinician panels is costly and slow, motivating the use of large language models (LLMs) as alternative adjudicators.

By Amy Rouillard, Sitwala Mundia, Linda Camara, Ziyaad Dangor, Michael Cameron Gramanie, Ismail Kalla, Shabir A. Madhi, Kajal Morar, Marlvin T. Ncube, Haroon Saloojee, Bruce A. Bassett