arXiv AI By Xi Fang, Weijie Xu, Yuchong Zhang, Stephanie Eckman, Scott Nickleach, Chandan K. Reddy

The Personalization Trap: How User Memory Alters Emotional Reasoning in LLMs

Read the original on arXiv AI →

arXiv:2510. 09905v2 Announce Type: replace Abstract: When an AI assistant remembers that Sarah is a single mother working two jobs, does it interpret her stress differently than if she were a wealthy executive?

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 18

To Memories and Beyond: From Remembering to Knowing You across Long-Term Multimodal Personal Archives

The paper introduces ReaLMem, a benchmark built from authentic multi‑year personal visual archives with first‑person annotations, designed to evaluate AI systems on factual recall, persona inference, and predictive personalization. It also proposes ChronoProfiler, a temporal‑weighting module that calculates stability scores for user attributes to resolve preference conflicts and enhance personalized decision making. Experiments with multimodal large language models and memory systems show that predictive personalization remains the hardest task, highlight performance gaps, and demonstrate that temporally informed representations significantly improve personalization.

By Wenqi Zhou, Zhuorui Yu, Kaiao Wen, Hao Zheng, Xinyi Zheng, Peiran Wu, Enmin Zhou, Chi-Hao Wu, Junxiao Shen
arXiv AI
Aug 19

Beyond BFI: The CSI for Enhanced Reliability and Validity in Evaluating LLM Personality Traits

The paper introduces the Core Sentiment Inventory (CSI), a new personality trait evaluation tool for large language models (LLMs) that addresses reliability and validity issues found in existing methods like the Big Five Inventory (BFI). CSI is designed specifically for LLMs, supports both English and Chinese, and provides detailed psychological portraits of model behavior. Experiments show that CSI captures nuanced behavioral patterns, improves reliability, and correlates strongly (above 0.85) with real-world LLM outputs.

By Huanhuan Ma, Haisong Gong, Xiaoyuan Yi, Xing Xie, Philip S. Yu, Dongkuan Xu