The study evaluates whether a fine‑tuned open‑weight model (Gemma‑3‑12B) can match the performance of GPT‑4o in extracting multi‑label intracranial hemorrhage acuity from non‑contrast head‑CT reports. Using a 2×2 design that varied adaptation strategy (classification head vs. instruction fine‑tuning) and training‑data source (distilled real GPT‑4o labels vs. synthetic GPT‑4o‑generated reports), the distilled instruction‑tuned model achieved macro‑F1 scores comparable to GPT‑4o and surpassed the untuned base model. The key finding is that the source of training data—distilled real reports—was more important than the fine‑tuning method, and that the entire fine‑tuning and inference process fits on a single 24 GB consumer GPU.
By Aawez Mansuri, Kush Mehta, Mohammadreza Chavoshi, Jahanzaib Malik, Theodorus Dapamede, Frank Li, Rohan Isaac, Beatrice Brown-Mulry, Chiratidzo Rudado Sanyika, YoungSeok Jeon, Judy W. Gichoya, Ali Emami, Hari Trivedi
arXiv:2606. 00986v1 Announce Type: new Abstract: Federated learning (FL) enables multiple data holders to train machine learning models collaboratively without centralizing raw data, making it useful in privacy sensitive domains such as healthcare and institutional data sharing.
By Ivo Osterberg Nilsson, Maximilian Birr Engvall, Viktor Valadi, Teddy Lazebnik
Converting free-text radiology reports into structured labels supports cohort building, quality assurance, and monitoring of clinical imaging models, but the strongest label extractors are hosted prop...
The paper introduces a clinically grounded privacy evaluation framework for medical language models, assessing leakage across a spectrum of adversarial access levels—from publicly inferable demographics to leaked note fragments. Using this framework on an LM pretrained on 378,000 clinical notes, the authors find that routine encounter metadata leads to high verbatim memorization and significant recovery of sensitive diagnoses (e.g., AUROC 0.91 for abortion, 0.82 for HIV). They also note that exact-match memorization can overstate disclosure, with 36% of memorized tokens being templated documentation, underscoring the risks of training on longitudinal clinical data and offering a reusable evaluation tool.
By Sasha Ronaghi, Sana Tonekaboni, Lena Stempfle, Vivian Utti, Jordan Li Cahoon, Nathaniel Hendrix, Ayin Vala, Marzyeh Ghassemi, Emily Alsentzer
Aegis is a client‑side defense for medical federated learning that protects against model inversion attacks by adding a masking gradient derived from locally synthesized data. The method exploits the fact that attacks fail when the effective batch size exceeds the model’s leakage capacity, turning this bottleneck into a privacy guarantee. Experiments on MNIST, CIFAR‑10, and MedMNIST datasets show that Aegis neutralizes state‑of‑the‑art attacks while preserving model accuracy and adding only modest overhead.
By Chaoyu Zhang, Shanghao Shi, Heng Jin, Ning Wang, Y. Thomas Hou, Wenjing Lou
arXiv:2606. 08769v1 Announce Type: cross Abstract: Automatic evaluation is critical for high-stakes text generation, where errors often involve omitted findings, hallucinated content, polarity reversals, location changes, uncertainty mismatches, and temporal-comparison errors rather than low surface similarity alone.
By Weixin Liu, Juming Xiong, Yang Li, Qingyuan Song, Susannah Rose, Murat Kantarcioglu, Bradley Malin, Zhijun Yin