arXiv AI By Ben Wang, Kang Zhou, Lifan Guo, Feng Chen, Chi Zhang

FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents

Read the original on arXiv AI →

arXiv:2608. 04095v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 4

The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis

The study examines how user context—such as memory, profiles, and role prompts—affects Large Language Models’ (LLMs) financial analysis. Using 3,575 SEC filings and twelve LLMs, the authors distinguish between evidence selection and interpretation, finding that most context spillover arises from differing interpretations under various roles rather than from retrieving different evidence. They evaluate two mitigation strategies—expressing investor mindset as a user profile instead of an assistant role, and separating evidence-based from personalized outputs—both of which reduce but do not eliminate spillover, with effectiveness varying across models.

By Ahmed Asaad, Amr Mohamed, Yang Zhang, Omneya Abdelsalam
Hugging Face Trending Papers
Sep 2

The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis

The paper investigates how user context—such as memory, profiles, and role prompts—affects large language models’ financial analysis. By testing 3,575 SEC filings across twelve LLMs, the study distinguishes between evidence selection and interpretation, finding that interpretation under different roles drives most user-context spillover. Two mitigation strategies—using a user profile instead of an assistant role and separating evidence-based from personalized outputs—reduce but do not eliminate this spillover, with effectiveness varying by model.

arXiv Machine Learning
Sep 3

CAPTURE: Disentangling Preference Drift from Memory Poisoning in Personalized LLM Agents

The paper introduces CAPTURE, a system designed to help personalized language agents distinguish genuine preference changes from temporary context shifts or malicious memory poisoning. CAPTURE employs a neural differential-equation belief tracker, a multi-timescale memory ledger, uncertainty-triggered clarification, and counterfactual auditing to resolve ambiguity. Experiments on 480 episodes from 96 users show CAPTURE outperforms baseline methods, limiting poisoning success while accepting most real preference updates.

By S M Asif Hossain, Ruksat Khan Shayoni, Md Kishor Morol
arXiv AI
Sep 3

AdaMem: Learning What to Remember with Adaptive Memory Policies for Personalized Agents

AdaMem introduces adaptive memory policies that allow personalized agents to decide what information to write into long‑term memory based on user preferences for each interaction context. Each policy is updated from periodic feedback and controls subsequent memory writing, aiming to improve relevance and reduce unnecessary memory persistence. In experiments on AdaMem‑Bench, AdaMem raises QA accuracy from 80.0% to 84.35% while cutting persistent memory by 9.27%, though models still struggle to execute policies reliably.

By Xingyu Chen, Rui Wang, Zhaopeng Tu, Liefeng Bo