arXiv Computation and Language

AI Models Can Predict and Collaboratively Modulate Human Memory Search

The study investigates how large language models (LLMs) can assist humans in semantic memory search tasks. By using the semantic fluency task (SFT), the researchers evaluate whether LLMs can follow and enhance human mental trajectories during generative semantic retrieval. Results show that an LLM’s ability to track and predict human memory trajectories in this task surpasses that of other humans.

arXiv AI
Aug 28

Artificial Intelligence Models Can Predict and Collaboratively Modulate Human Memory Search

The article reports that large language models can predict and collaboratively modulate human memory search during a semantic fluency task. By tracking and forecasting participants’ semantic retrieval patterns, the models outperform other humans in following these mental trajectories. This suggests that AI can serve as a cognitive tool to extend human abilities in open‑ended conceptual exploration and creative ideation.

By Eric Lacosse, Mariana Duarte, Graham Todd, Peter M. Todd, Daniel C. McNamee
arXiv AI
Sep 17

WordPolo: Evaluating Language Models Through Iterative Semantic Feedback

WordPolo is a word‑finding task that evaluates language models by having them guess an unknown target word and receive semantic similarity feedback. Participants start with no knowledge, make iterative guesses, and receive distance scores that guide them through semantic space. The study tests recent LLMs, LRMs, humans, and a heuristic on 1,500 puzzles, revealing that while solve rates vary widely, many models make meaningful progress and exhibit human‑like strategies, highlighting the importance of assessing reasoning processes, not just final accuracy.

By Tyler McDonald, Ali Emami
arXiv AI
Aug 18

AutoMem: A Text-Gradient Recursive Self-Improvement Framework for Automated Memory Architectures Search

arXiv:2608. 14621v1 Announce Type: cross Abstract: Long-term memory is increasingly central to LLM agents, yet memory design remains a highly coupled architecture problem: what to encode, how to store it, how to retrieve it, and how to manage it can vary substantially across tasks and backbone models.

By Lin Du, Jie Zhou, Yuxuan Cai, Kai Chen, Qin Chen, Xin Li, Bo Zhang, Wei Li, Liang He
arXiv AI
Sep 4

When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents

The paper introduces LOCOMO-CONV, a conversational memory benchmark that expands on the existing LoCoMo dataset with four query styles—dialog, implicit, counterfactual, and composed—designed to evaluate memory systems in realistic conversational settings. Experiments across five memory systems reveal that conversational framing uncovers significant retrieval gaps missed by traditional QA benchmarks, particularly for implicit and composed queries, and that strong retrieval does not necessarily translate into higher response quality. The study also highlights silent grounding in implicit queries, where memory enhances contextual grounding without explicitly presenting the gold fact, suggesting a need for reasoning-based memory elaboration.

By Wen-Yu Chang, Yun-Nung Chen