APeB: Benchmarking Personalization Ability of Large Language Model Agents
arXiv:2607. 03162v1 Announce Type: new Abstract: LLM-powered agents struggle with personalization when users issue raw, underspecified queries.
arXiv:2607. 20482v1 Announce Type: new Abstract: Recent advances in large language models have enabled web agents to autonomously execute complex tasks.
arXiv:2607. 03162v1 Announce Type: new Abstract: LLM-powered agents struggle with personalization when users issue raw, underspecified queries.
arXiv:2511. 12997v2 Announce Type: replace Abstract: Multimodal LLM-powered agents have recently demonstrated impressive capabilities in web navigation, enabling agents to complete complex browsing tasks across diverse domains.
MemoryArena is a new evaluation gym that benchmarks agent memory in interdependent multi‑session tasks. Unlike prior benchmarks that test memorization or single‑session action in isolation, MemoryArena requires agents to acquire memory while interacting with the environment and then use that memory to guide future decisions across a range of tasks such as web navigation, planning, information search, and formal reasoning. The benchmark reveals that agents excelling on existing long‑context memory tests perform poorly here, highlighting a gap in current memory evaluation methods.
The paper introduces the Multi-Session Personalized Tool Calling (MPT) benchmark, containing 4,695 instances across 459 multi‑session histories that test Preference Recall, Induction, and Transfer. It proposes PRefine, a test‑time memory method that refines a user’s latent preference via a generate‑verify‑refine loop. Experiments with five LLMs show that PRefine outperforms existing memory systems and even full‑history prompting on Preference Transfer, suggesting that personalized agents should encode behavior as preferences rather than merely storing past interactions.
arXiv:2609.36488v1 Announce Type: new Abstract: Large language model (LLM) agents have demonstrated strong performance on complex web navigation tasks, yet they remain brittle in real-world settings...
arXiv:2607. 27056v1 Announce Type: new Abstract: Personalized agents are increasingly applied to assist users across a wide range of tasks.
The paper introduces HiPS, a hierarchical strategy co‑evolution framework for memory‑augmented agents that separates memory management into a globally shared foundation and a user‑specific adaptive tier. HiPS uses a Universal Strategy to capture shared principles from cross‑persona trajectories, Persona Delta Distillation to create tailored rules for users deviating from general patterns, and Cross‑Level Rule Flow to dynamically adjust the boundary between global and personal rules. Experiments show that this approach consistently outperforms existing memory‑augmented baselines.
arXiv:2609.08599v1 Announce Type: new Abstract: Large Language Model (LLM) agents are evolving from single-session tools toward long-term personal assistants that must adapt to individual users acros...
ReMem is a new recommendation agent framework that rethinks perception and memory for long-context recommendation tasks. It replaces raw HTML parsing with OCR-based multimodal perception from screenshots, extracting structured information in a platform-agnostic way. The framework also introduces a chunk-wise sequential memory update strategy and a multi-memory GRPO variant to efficiently model evolving user preferences over arbitrarily long interaction histories, achieving a 5.16% average improvement over state-of-the-art baselines on three recommendation agent tasks.
Large Language Model (LLM) agents are evolving from single-session tools toward long-term personal assistants that must adapt to individual users across tasks, contexts, and interactions. This shift m...
arXiv:2506. 01952v2 Announce Type: replace-cross Abstract: Powered by large language models (LLMs), web browsing agents operate graphical user interfaces in a human-like manner, offering a transparent and general framework for automating web-based tasks.
arXiv:2609.24971v1 Announce Type: new Abstract: Agents today often take real-world actions that depend on long-term memory and context recall over time. However, most current memory benchmarks are bu...