arXiv AI By Dannong Wang, Yuran Zhang, Bian Sun, Alex Stinard, Yuzhang Shang, Song Wang, Yu Tian

PrivMeSA: Privacy-Aware Self-Evolving Multi-Agent System for Medicine via Local-Remote LLM Collaboration

Read the original on arXiv AI →

PrivMeSA is a privacy‑aware, self‑evolving multi‑agent system that enables local clinical LLM agents to consult remote specialists while controlling patient data disclosure. It uses reinforcement learning to balance task accuracy with privacy risk, evaluates risk over entire outbound transcripts, and employs a local lesson memory to distill and reuse expertise without further remote exchanges. On an emergency‑department benchmark, PrivMeSA boosts task accuracy by up to 15.8 percentage points and virtually eliminates patient re‑identification risk, reducing personal detail disclosure from 98.0% to 0.2% of cases.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 1

CARE: Privacy-Compliant Agentic Reasoning with Evidence Discordance

The paper introduces MIMIC-DOS, a dataset derived from MIMIC-IV that focuses on ICU cases where patient symptoms and medical signs are discordant. It presents CARE, a privacy‑compliant multi‑stage agentic reasoning framework that uses a proprietary LLM to generate structured categories and transitions, while a local LLM performs evidence acquisition and decision‑making. In retrospective evaluations on MIMIC‑DOS, CARE outperforms other LLMs and agentic workflows, demonstrating stronger handling of conflicting clinical evidence while preserving patient privacy.

By Haochen Liu, Weien Li, Rui Song, Zeyu Li, Chun Jason Xue, Xiao-Yang Liu, Sam Nallaperuma-Herzberg, Xue Liu, Ye Yuan
arXiv Computation and Language
Aug 25

Clinically Grounded Privacy Evaluation of Medical LMs

The paper introduces a clinically grounded privacy evaluation framework for medical language models, assessing leakage across a spectrum of adversarial access levels—from publicly inferable demographics to leaked note fragments. Using this framework on an LM pretrained on 378,000 clinical notes, the authors find that routine encounter metadata leads to high verbatim memorization and significant recovery of sensitive diagnoses (e.g., AUROC 0.91 for abortion, 0.82 for HIV). They also note that exact-match memorization can overstate disclosure, with 36% of memorized tokens being templated documentation, underscoring the risks of training on longitudinal clinical data and offering a reusable evaluation tool.

By Sasha Ronaghi, Sana Tonekaboni, Lena Stempfle, Vivian Utti, Jordan Li Cahoon, Nathaniel Hendrix, Ayin Vala, Marzyeh Ghassemi, Emily Alsentzer
arXiv Machine Learning
Sep 16

Memorisation bias in medical AI

arXiv:2609.17223v1 Announce Type: new Abstract: Medical AI models hold immense potential to improve patient outcomes, but they are also known to unintentionally memorise individual records from their...

By Moritz A. Knolle, Martin J. Menten, Laurin Lux, M\'elanie Roschewitz, Emma A. M. Stanley, Georgios Kaissis, Daniel Rueckert, Ben Glocker