arXiv AI By Saurabh Ranjan, Konstantina Sokratous, Brian Odegaard

Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory

Read the original on arXiv AI →

arXiv:2607. 23927v1 Announce Type: new Abstract: A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 1

UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory

UTILMEM is a new diagnostic benchmark that tests how conversational agents use long‑term memory, focusing on reasoning over dense histories, spotting implicitly relevant memories, synthesizing distributed evidence, and resisting interference from similar distractors. It contains 1,717 instances across five domains and evaluates a range of retrieval‑based and memory‑augmented systems. The study shows that strong performance on traditional factual recall does not guarantee effective memory utilization, highlighting a gap between retrieving information and integrating it into coherent, task‑oriented outputs.

By Peijun Qing, Fobo Shi, Soroush Vosoughi
arXiv AI
2d ago

RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations

RealCompanion is a benchmark that evaluates an AI companion’s ability to understand a human over long, real-world conversations. It consists of ten real relationships with 27,218 messages spanning up to 120 days, along with derived files such as a profile, persona, chat ground truth, and a question set that cites the relevant messages. The study finds that past context is rarely needed, memory detection is challenging, and agent systems vary widely in cost while achieving similar persona reconstruction.

By Arman Behnam, Sunglyoung Kim, Liangwei Yang
arXiv AI
Sep 17

MIRAGE: How Conversation State Shapes Historical Evidence Use in Multimodal Personal Agents

MIRAGE is a controlled study that examines how multimodal personal agents use historical evidence when conversation state changes. The study keeps evidence, questions, and scoring constant while varying only the conversation state, then checks if agents can determine answerability, recover the correct source, and answer from it. Results across seven multimodal backbones show distinct failure regimes before and after compaction, heavy reliance on context continuity by open-weight models, and mixed effects of retrieval pressure on source attribution.

By Yu Liu, Wenxiao Zhang, Cheng Hu, Cong Cao, Fangfang Yuan, Xinyu Wang, Jin B. Hong, Yanbing Liu