arXiv AI By Arman Behnam, Sunglyoung Kim, Liangwei Yang

RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations

Read the original on arXiv AI →

RealCompanion is a benchmark that evaluates an AI companion’s ability to understand a human over long, real-world conversations. It consists of ten real relationships with 27,218 messages spanning up to 120 days, along with derived files such as a profile, persona, chat ground truth, and a question set that cites the relevant messages. The study finds that past context is rarely needed, memory detection is challenging, and agent systems vary widely in cost while achieving similar persona reconstruction.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 1

UTILMEM: Benchmarking Evidence Utilization in Long-Term Conversational Memory

UTILMEM is a new diagnostic benchmark that tests how conversational agents use long‑term memory, focusing on reasoning over dense histories, spotting implicitly relevant memories, synthesizing distributed evidence, and resisting interference from similar distractors. It contains 1,717 instances across five domains and evaluates a range of retrieval‑based and memory‑augmented systems. The study shows that strong performance on traditional factual recall does not guarantee effective memory utilization, highlighting a gap between retrieving information and integrating it into coherent, task‑oriented outputs.

By Peijun Qing, Fobo Shi, Soroush Vosoughi
arXiv Computer Vision
Sep 7

ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory

ICM-Bench is a new benchmark for evaluating identity-centric reasoning in multimodal agents with long-term memory. It consists of 839 synthetic video clips totaling 141 minutes and 1,217 open-ended questions about six recurring adults in a one-year life album. The benchmark isolates the ability to maintain recurring person identities and reason over their cross-time relations, and compares various baseline systems, showing that while Gemini 3.1 Pro performs well overall, its accuracy drops on questions requiring long-term identity profiles.

By Shidu Ren, Yunze Liu, Xing Liu, Chi-Hao Wu, Enmin Zhou, Junxiao Shen