arXiv AI

PrivMeSA: Privacy-Aware Self-Evolving Multi-Agent System for Medicine via Local-Remote LLM Collaboration

PrivMeSA is a privacy‑aware, self‑evolving multi‑agent system that enables local clinical LLM agents to consult remote specialists while controlling patient data disclosure. It uses reinforcement learning to balance task accuracy with privacy risk, evaluates risk over entire outbound transcripts, and employs a local lesson memory to distill and reuse expertise without further remote exchanges. On an emergency‑department benchmark, PrivMeSA boosts task accuracy by up to 15.8 percentage points and virtually eliminates patient re‑identification risk, reducing personal detail disclosure from 98.0% to 0.2% of cases.

arXiv Computation and Language
Sep 1

CARE: Privacy-Compliant Agentic Reasoning with Evidence Discordance

The paper introduces MIMIC-DOS, a dataset derived from MIMIC-IV that focuses on ICU cases where patient symptoms and medical signs are discordant. It presents CARE, a privacy‑compliant multi‑stage agentic reasoning framework that uses a proprietary LLM to generate structured categories and transitions, while a local LLM performs evidence acquisition and decision‑making. In retrospective evaluations on MIMIC‑DOS, CARE outperforms other LLMs and agentic workflows, demonstrating stronger handling of conflicting clinical evidence while preserving patient privacy.

By Haochen Liu, Weien Li, Rui Song, Zeyu Li, Chun Jason Xue, Xiao-Yang Liu, Sam Nallaperuma-Herzberg, Xue Liu, Ye Yuan
arXiv Computation and Language
Aug 25

Clinically Grounded Privacy Evaluation of Medical LMs

The paper introduces a clinically grounded privacy evaluation framework for medical language models, assessing leakage across a spectrum of adversarial access levels—from publicly inferable demographics to leaked note fragments. Using this framework on an LM pretrained on 378,000 clinical notes, the authors find that routine encounter metadata leads to high verbatim memorization and significant recovery of sensitive diagnoses (e.g., AUROC 0.91 for abortion, 0.82 for HIV). They also note that exact-match memorization can overstate disclosure, with 36% of memorized tokens being templated documentation, underscoring the risks of training on longitudinal clinical data and offering a reusable evaluation tool.

By Sasha Ronaghi, Sana Tonekaboni, Lena Stempfle, Vivian Utti, Jordan Li Cahoon, Nathaniel Hendrix, Ayin Vala, Marzyeh Ghassemi, Emily Alsentzer
arXiv Machine Learning
Sep 16

Memorisation bias in medical AI

arXiv:2609.17223v1 Announce Type: new Abstract: Medical AI models hold immense potential to improve patient outcomes, but they are also known to unintentionally memorise individual records from their...

By Moritz A. Knolle, Martin J. Menten, Laurin Lux, M\'elanie Roschewitz, Emma A. M. Stanley, Georgios Kaissis, Daniel Rueckert, Ben Glocker
Hugging Face Trending Papers
Jun 8

Experience Makes Skillful: Enabling Generalizable Medical Agent Reasoning via Self-Evolving Skill Memory

Medical agent systems are increasingly expected to support interactive clinical decision making rather than only static question answering. In such settings, effective agents must reuse prior experience across evolving cases, yet existing memory mechanisms often retain raw historical traces that are redundant, noisy, and difficult to govern.

arXiv Machine Learning
Sep 14

AIM: A Privacy-Aware Interoperable Memory Framework for Multi-Agent Multi-User LLM Systems

The paper introduces AIM, a privacy‑aware memory framework that lets multi‑agent, multi‑user large language models manage both private and shared memory. AIM classifies data as private (user‑specific) or public (shared) and enforces index‑level access controls to protect sensitive information while enabling shared knowledge to improve coordination. The authors also present MUMBench, a new dataset for evaluating memory operations in multi‑user settings, and report high accuracy metrics for AIM on this benchmark.

By Zachary Johnson, Nigel Boachie Kumankumah, Somya Chatterjee, Tejas Sathyamurthi, Min Chen, Xinyi Alice Li, Xiao Wang, Emily Morgan Gelchie, Jessica Lin, Sadid A. Hasan, Sulaiman Vesal
arXiv AI
Aug 25

MACD: Multi-Agent Clinical Diagnosis with Self-Learned Knowledge for LLM

The paper introduces MACD, a Multi-Agent Clinical Diagnosis framework that enables large language models to self‑learn clinical knowledge through a multi‑agent pipeline of summarization, refinement, and application. MACD is extended into a human‑AI collaborative workflow where multiple diagnostician agents consult iteratively, guided by a judge agent and human oversight. Evaluation on the MIMIC‑MACD cohort shows significant gains in diagnostic accuracy—an average 11.6 percentage‑point improvement over authoritative knowledge for open‑weight LLMs and an 18.3‑percentage‑point boost over physician‑only diagnosis in text‑only vignettes.

By Wenliang Li, Rui Yan, Xu Zhang, Li Chen, Hongji Zhu, Jing Zhao, Junjun Li, Mengru Li, Wei Cao, Zihang Jiang, Wei Wei, Kun Zhang, Shaohua Kevin Zhou
arXiv AI
Aug 10

ResidencyRL: Reinforcement Learning in Simulated Clinical Environments

arXiv:2608. 07418v1 Announce Type: new Abstract: In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters, with diverse sources of feedback and progressively greater autonomy.

By Valentin Li\'{e}vin, Samuel Schmidgall, Tim Strother, Alex Bijamov, Akshay Goel, Anil Palepu, Chunjong Park, Vahid Balazadeh, Min Woo Sun, Marius Guerard, Justin Chen, Dave Steiner, Vikram Dhillon, Ibrahim Azar, Akhil Mehta, Nicholas Spetsieris, Shilpan Shah, Maen Abdelrahim, Amit Dahiya, Yun Liu, Katherine Chou, Yossi Matias, Avinatan Hassidim, Dale R. Webster, Quoc V. Le, Raia Hadsell, Joelle Barral, Carey Radebaugh, Aleksandra Faust, Shekoofeh Azizi, Mike Schaekermann, Po-Hsuan Cameron Chen, Tao Tu, David Racz, Lin Yang