Adopt $\neq$ Adapt: Longitudinal Analyses of LLM Conversations in the Wild
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2608. 05246v1 Announce Type: new Abstract: Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities.
arXiv:2609.01257v1 Announce Type: new Abstract: As LLM-based human simulators are increasingly used for policy, evaluation, and training, they must faithfully reproduce real behavioral patterns. Whil...
arXiv:2609.07594v1 Announce Type: cross Abstract: Conversational Task Assistants (CTAs) are multimodal dialogue systems that support users in complex real-world tasks such as cooking and DIY through...
arXiv:2607. 27056v1 Announce Type: new Abstract: Personalized agents are increasingly applied to assist users across a wide range of tasks.
PersonaMem-v3 is a benchmark and evaluation harness designed to assess omni-platform personal intelligence for AI agents. It is built from over one million anonymized real-world engagement histories, covering social media, chatbots, calendars, and AI companions, and tracks user preferences and habits over time. The benchmark tests agents on personalization, LLM-powered recommendation, proactiveness, agentic tool use, and geo-temporal reasoning, evaluating their ability to infer holistic user understanding, personalize responses, rerank recommendations, follow user steering, and avoid inappropriate personalization.
arXiv:2607. 20734v1 Announce Type: new Abstract: As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction.