CliniCIRCA is a modular large‑language‑model framework that reconstructs longitudinal mental‑health patient journeys from raw electronic health record narratives. It temporally classifies clinical events in unstructured discharge summaries without explicit timestamps, producing 15,891 tagged events from 52 summaries and correcting 629 errors to create verified gold‑standard timelines. The framework then generates temporally grounded summaries, compressing each source by 1.52×, and scales to produce 1,000 silver‑standard timelines for training, showing that instruction tuning improves event extraction, temporal tagging, and summarization across models.
By Aiwei Ivy Zhang, Nimra Ishfaq, Mohit Chandra, Santiago Alvarez Lesmes, Adam Coscia, Khatiya Chelidze Moon, Xiaohan Ding, Munmun De Choudhury
The paper "Beyond Static Summarization: Proactive Memory Extraction for LLM Agents" identifies two shortcomings in current memory extraction for large language model agents: (1) extraction occurs ahead of time and mixes multiple types of information, leading to loss of useful details, and (2) extraction is typically one‑off, allowing errors and hallucinations to persist. To address these issues, the authors propose ProMem, a proactive framework that separates details, events, and relations, applies distinct extraction strategies for each, checks for completeness, and verifies facts at an atomic level. Experiments demonstrate that ProMem enhances memory completeness and question‑answering accuracy while maintaining a favorable balance between quality and token cost.
By Chengyuan Yang, Zequn Sun, Wei Wei, Wei Hu
OmniConfess is a training‑free method designed to reduce hallucinations in omni‑modal large language models (OmniLLMs) that handle text, images, audio, and video. The approach fixes a candidate response and re‑scores it at token resolution while selectively intervening on evidence from each modality, producing a token‑by‑channel confession that shows which evidence supports each part of the response. Using this confession, OmniConfess preserves grounded content and corrects commitments that rely on irrelevant or contradictory evidence. The authors evaluated the method on OmniHalluBench, a 3,540‑example benchmark drawn from six datasets across multiple modalities and tasks, and found that OmniConfess mitigates hallucinations across diverse settings.
By Huiqiang Rong, Haoran Luo, Hui Feng, Zhonghong Ou, Kaiwen Xue, Guoxin Zhang, Yifan Zhu
arXiv:2605. 28910v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown promise on summarization tasks, but they often produce hallucinations, which are unsupported or incorrect statements that limit their reliability in specialized healthcare applications.
By Shamanth Kuthpadi Seethakantha, Dung Ngoc Thai, Vara Prasad Gudi, Simran Tiwari, Rami Matar, Avijit Mitra, Wenlong Zhao, Andrew McCallum, Wael Salloum
The paper introduces DETECT-REMASK-REPAIR, a diffusion-based method for updating outdated spans in existing summaries while keeping supported content intact. It identifies, masks, and repairs only the changed regions using masked diffusion language models. Experiments on DialogSum and a new StreamSum benchmark show that this localized repair improves faithfulness, reduces repair time to under half a second, and offers trade‑offs between faithfulness, speed, and preservation of the original summary.
By Hao Zou, Zachary Horvitz, Chandhru Karthick, Zhou Yu, Kathleen McKeown
arXiv:2512. 21577v3 Announce Type: replace-cross Abstract: Despite numerous attempts at mitigation since the inception of language models, hallucinations remain a persistent problem even in today's frontier LLMs.
By Emmy Liu, Varun Gangal, Chelsea Zou, Michael Yu, Xiaoqi Huang, Alex Chang, Zhuofu Tao, Karan Singh, Sachin Kumar, Steven Y. Feng