arXiv Computation and Language

FTA-Mem: Fact-Time-Affect Anchored Memory for Low-Density Long-Term Dialogue

arXiv Computation and Language
Sep 14

CueMem: Cue-Guided Context Reconstruction for Long-Term Conversational Memory

CueMem is a cue‑guided framework for long‑term conversational memory that reconstructs query‑relevant dialogue context from compressed memory records. Instead of treating memory units as self‑contained evidence, it extracts fine‑grained cues linked to their source turns and, at query time, expands from these cues over a turn graph to rebuild a compact evidence context. Experiments on LoCoMo and LongMemEval show that CueMem outperforms baseline memory methods, reduces input tokens and latency, and improves long‑term conversational question answering.

By Changjian Wang, Rongzhen Li, Weili Guan, Shuming Shi, Quan Lu, Ning Jiang
arXiv AI
Sep 4

When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents

The paper introduces LOCOMO-CONV, a conversational memory benchmark that expands on the existing LoCoMo dataset with four query styles—dialog, implicit, counterfactual, and composed—designed to evaluate memory systems in realistic conversational settings. Experiments across five memory systems reveal that conversational framing uncovers significant retrieval gaps missed by traditional QA benchmarks, particularly for implicit and composed queries, and that strong retrieval does not necessarily translate into higher response quality. The study also highlights silent grounding in implicit queries, where memory enhances contextual grounding without explicitly presenting the gold fact, suggesting a need for reasoning-based memory elaboration.

By Wen-Yu Chang, Yun-Nung Chen
arXiv Computation and Language
Sep 1

UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory

UTILMEM is a new diagnostic benchmark that tests how conversational agents use long‑term memory, focusing on reasoning over dense histories, spotting implicitly relevant memories, synthesizing distributed evidence, and resisting interference from similar distractors. It contains 1,717 instances across five domains and evaluates a range of retrieval‑based and memory‑augmented systems. The study shows that strong performance on traditional factual recall does not guarantee effective memory utilization, highlighting a gap between retrieving information and integrating it into coherent, task‑oriented outputs.

By Peijun Qing, Fobo Shi, Soroush Vosoughi
arXiv Computation and Language
Sep 15

Where to Look and What to Use: Retrieve-Localize-Generate for Long-Term Conversational Memory Question Answering

arXiv:2609.07093v2 Announce Type: replace Abstract: Retrieval-augmented generation (RAG) enables large language models (LLMs) to answer questions by accessing external knowledge and has been widely a...

By Yifan Wang, Xinkui Lin, Yongxiu Xu, Shen Gao, Ruochen Yang, Kun Huang, Yubin Wang, Jie Wu, Wei Liu, Jian Luan, Hongbo Xu, Shuo Shang
arXiv AI
4d ago

APDMem: Agent-Controlled Progressive Disclosure for Query-Adaptive Long-Term Memory

APDMem is a hierarchical long‑term memory system for personalized LLM assistants that uses progressive disclosure to retrieve conversation history. It organizes memory into four layers—thematic summaries, personalized key facts, turn‑level evidence notes, and raw messages—and a controller reads from the top layer, drilling down only when necessary. This approach balances cost and fidelity, allowing simple queries to finish early while deeper inspection is triggered for complex or exact‑evidence requests, and a note synthesizer structures retrieved evidence before answer generation. Experiments on LongMemEval show APDMem performs strongly while accessing only 8% of the total conversations.

By Chin-Lun Fu, Anagha Kulkarni, Hong Ni, Behrouz Madahian
arXiv AI
Sep 3

Beyond Dialogue Time: Temporal Semantic Memory for Personalized LLM Agents

The paper introduces Temporal Semantic Memory (TSM), a framework that improves how large language model agents manage memory by addressing two key shortcomings: temporal inaccuracy and temporal fragmentation. TSM constructs a semantic timeline instead of a dialogue timeline, consolidating temporally continuous and semantically related information into durative memory. During retrieval, it aligns the query’s temporal intent with the semantic timeline, enabling the use of temporally appropriate durative memories and yielding up to a 12.2% accuracy boost over existing methods.

By Miao Su, Yucan Guo, Zhongni Hou, Long Bai, Zixuan Li, Yufei Zhang, Guojun Yin, Wei Lin, Xiaolong Jin, Jiafeng Guo, Xueqi Cheng
Hugging Face Trending Papers
Aug 2

TrajWiki: Source-Grounded Memory Trajectories for Long-Horizon Dialogue Agents

Large language model agents have shown strong capabilities in generating coherent and contextually appropriate responses, yet robust long-horizon dialogue remains limited by the lack of external memory that is traceable, updatable, and diagnostically transparent. Existing memory-augmented agents often store memories as isolated records or overwritable states, making it difficult to preserve how information originates, evolves, conflicts, or becomes obsolete over time.

arXiv AI
Jun 2

Connecting the Dots: Benchmarking Reflective Memory in Long-Horizon Dialogue

arXiv:2606. 01223v1 Announce Type: cross Abstract: Despite substantial progress in long-context modeling, existing benchmarks remain confined to factual memory for explicit recall, failing to measure the reflective memory required to synthesize fragmented, multimodal cues into high-level interpretations.

By Jingjie Lin, Bingbing Wang, Zihan Wang, Zhengda Jin, Weiming Qiao, Jing Li, Ruifeng Xu
arXiv AI
Oct 2

CAVE-Mem: Boundary-Aware Experience Validation for Memory Search

CAVE-Mem is a training‑free framework that enhances memory search for long‑term memory agents by treating experience as a typed intervention operator with conditions on applicability, boundary, and utility. It first retrieves a base answer and then only applies an intervention if the operator matches the current memory substrate, answer contract, evidence boundary, and cross‑fitted utility; otherwise it abstains. Experiments on conversational memory, multi‑hop QA, and long‑document reasoning demonstrate consistent improvements over relevance‑only experience reuse.

By Xinyu Li