arXiv Computation and Language

ReLay: Personalized LLM-Generated Plain-Language Summaries for Better Understanding, but at What Cost?

arXiv AI
Jun 9

Personalization Meets Safety:Mechanisms,Risks,and Mitigations in Personalized LLMs

arXiv:2606. 09038v1 Announce Type: new Abstract: Large Language Models (LLMs) have enabled increasingly personalized interactions by adapting to users' preferences, contexts, and long-term histories.

By Yanyan Luo, Xue Han, Ruiqiao Bai, Xin Huang, Yitong Wang, Qian Hu, Qing Wang, Chunxu Zhao, Jie Liu, Cong Geng, Lehao Xing, Pengwei Hu, Junlan Feng
arXiv AI
Jun 6

Evaluating the Utility of Personal Health Records in Personalized Health AI

arXiv:2605. 18937v2 Announce Type: replace Abstract: Patient-managed Personal Health Records (PHRs) promises to empower patients to better understand their health; but information in the record is complex, potentially hindering insights.

By Rory Sayres, Kejia Chen, Ayush Jain, Matthew Thompson, Jonathan Richina, Xiang Yin, Jimmy Hu, Fan Zhang, Bob Lou, Mike Sanchez, Ines Mezerreg, Meredith Schreier, Hamsa Subramaniam, I-Ching Lee, Yugang Jia, Daniel Mcduff, Yossi Matias, Avinatan Hassidim, Dale Webster, Yun Liu, Jackie Barr, Quang Duong
arXiv AI
Aug 28

Evaluating AI Generated Summaries for Cancer Patients

The study evaluates AI-generated summaries for cancer patients using a dual assessment framework that includes human experts and LLM-as-a-judge. Human domain experts—oncology clinicians and patient-facing care staff—assess summary quality on accuracy, clinical relevance, and readability. The research identifies limitations such as omissions and minor inaccuracies, which are then used to iteratively refine prompts, grounding, and safety guardrails.

By Muhammad Aurangzeb Ahmad, Kim Shyu, Leon Oliver, Fergus Sleight, Paul Landau
arXiv Computation and Language
Sep 3

Improving Health Literacy through Lay Summarization of Radiological Reports: An Evaluation of BioNER and Retrieval-Augmented Generation

The paper examines how Retrieval-Augmented Generation (RAG) and Named Entity Recognition (NER) affect the quality of lay summaries of radiology reports. Using a framework that extracts clinically relevant findings via NER and grounds them with RAG, the authors evaluate few‑shot and fine‑tuned versions of Qwen and BioBART. Results show that NER consistently improves readability and overall quality, RAG alone offers no benefit and can introduce hallucinations, and the best performance comes from fine‑tuned BioBART with NER.

By Egecan \c{C}elik Evgin, \.Ilknur Karadeniz, Olcay Taner Y{\i}ld{\i}z
arXiv AI
3d ago

Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts

The paper introduces a benchmark of over 6,000 clinical triage scenarios, 7,000 physician annotations, and 225,000 large language model (LLM) responses to assess how LLMs perform under realistic variations in clinical text. The study finds that LLMs tend to recommend unnecessary care more often than physicians, especially when the input text is perturbed, and that LLM recommendations are more sensitive to gender and tone changes than human recommendations. These findings underscore the importance of deployment‑oriented evaluations that reflect expert physician behavior.

By Abinitha Gourabathina, Haoran Zhang, Yuexing Hao, Walter Gerych, Marzyeh Ghassemi
arXiv Machine Learning
Aug 12

Automated Data Enrichment using Confidence-Aware Fine-Grained Debate among Open-Source LLMs for Mental Health and Online Safety

arXiv:2512. 06227v3 Announce Type: replace-cross Abstract: Real-world indicators play an important role in many Natural Language Processing (NLP) applications, such as life events for mental health analysis and risky behaviours for online safety, yet labelling such information is often costly and/or difficult due to its multi-label and dynamic nature.

By Junyu Mao, Anthony Hills, Talia Tseriotou, Maria Liakata, Aya Shamir, Dan Sayda, Dana Atzil-Slonim, Natalie Djohari, Pamela Ugwudike, Mahesan Niranjan, Stuart E. Middleton