arXiv AI

NTS-CoT: Mitigating Hallucinations in LLM-based News Timeline Summarization with Chain-of-Thought Reasoning

arXiv:2606. 13171v1 Announce Type: cross Abstract: The rapid updates of online news make tracking event developments challenging, highlighting the need for timeline summarization (TLS).

arXiv AI
Sep 18

CliniCIRCA: A Modular LLM Framework for Constructing Longitudinal Mental Health Patient Journeys from Raw EHR Narratives

CliniCIRCA is a modular large‑language‑model framework that reconstructs longitudinal mental‑health patient journeys from raw electronic health record narratives. It temporally classifies clinical events in unstructured discharge summaries without explicit timestamps, producing 15,891 tagged events from 52 summaries and correcting 629 errors to create verified gold‑standard timelines. The framework then generates temporally grounded summaries, compressing each source by 1.52×, and scales to produce 1,000 silver‑standard timelines for training, showing that instruction tuning improves event extraction, temporal tagging, and summarization across models.

By Aiwei Ivy Zhang, Nimra Ishfaq, Mohit Chandra, Santiago Alvarez Lesmes, Adam Coscia, Khatiya Chelidze Moon, Xiaohan Ding, Munmun De Choudhury
arXiv AI
Sep 2

Beyond Static Summarization: Proactive Memory Extraction for LLM Agents

The paper "Beyond Static Summarization: Proactive Memory Extraction for LLM Agents" identifies two shortcomings in current memory extraction for large language model agents: (1) extraction occurs ahead of time and mixes multiple types of information, leading to loss of useful details, and (2) extraction is typically one‑off, allowing errors and hallucinations to persist. To address these issues, the authors propose ProMem, a proactive framework that separates details, events, and relations, applies distinct extraction strategies for each, checks for completeness, and verifies facts at an atomic level. Experiments demonstrate that ProMem enhances memory completeness and question‑answering accuracy while maintaining a favorable balance between quality and token cost.

By Chengyuan Yang, Zequn Sun, Wei Wei, Wei Hu
arXiv Computation and Language
1d ago

OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination

OmniConfess is a training‑free method designed to reduce hallucinations in omni‑modal large language models (OmniLLMs) that handle text, images, audio, and video. The approach fixes a candidate response and re‑scores it at token resolution while selectively intervening on evidence from each modality, producing a token‑by‑channel confession that shows which evidence supports each part of the response. Using this confession, OmniConfess preserves grounded content and corrects commitments that rely on irrelevant or contradictory evidence. The authors evaluated the method on OmniHalluBench, a 3,540‑example benchmark drawn from six datasets across multiple modalities and tasks, and found that OmniConfess mitigates hallucinations across diverse settings.

By Huiqiang Rong, Haoran Luo, Hui Feng, Zhonghong Ou, Kaiwen Xue, Guoxin Zhang, Yifan Zhu
arXiv AI
Jun 2

Hallucination Detection-Guided Preference Optimization for Clinical Summarization

arXiv:2605. 28910v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have shown promise on summarization tasks, but they often produce hallucinations, which are unsupported or incorrect statements that limit their reliability in specialized healthcare applications.

By Shamanth Kuthpadi Seethakantha, Dung Ngoc Thai, Vara Prasad Gudi, Simran Tiwari, Rami Matar, Avijit Mitra, Wenlong Zhao, Andrew McCallum, Wael Salloum
arXiv Computation and Language
Sep 7

Detect, Remask, Repair: Diffusion Editing for Faithful Summarization of Evolving Contexts

The paper introduces DETECT-REMASK-REPAIR, a diffusion-based method for updating outdated spans in existing summaries while keeping supported content intact. It identifies, masks, and repairs only the changed regions using masked diffusion language models. Experiments on DialogSum and a new StreamSum benchmark show that this localized repair improves faithfulness, reduces repair time to under half a second, and offers trade‑offs between faithfulness, speed, and preservation of the original summary.

By Hao Zou, Zachary Horvitz, Chandhru Karthick, Zhou Yu, Kathleen McKeown
arXiv Machine Learning
Sep 15

Omni-Streaming Thinking

arXiv:2609.15128v1 Announce Type: new Abstract: Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support...

By Enjun Du, Siyi Liu, Ziyu Zheng, Jingyu Li, Yiwen Guo, Yongqi Zhang, Difan Zou
arXiv Computation and Language
Aug 21

When Text and Numbers Disagree: Evidence Arbitration in Large Language Models

arXiv:2608. 20116v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence.

By Mattia Carletti, Edward Phillips, Fredrik K. Gustafsson, Patitapaban Palo, Lei Clifton, Danielle Belgrave, Xiao Gu, David A. Clifton