arXiv AI

LLMs as Oracles: Reliance on LLMs for Subjective Personal Questions

arXiv AI
Sep 4

User Perceptions vs. Proxy LLM Judges: Privacy and Helpfulness in LLM Responses to Privacy-Sensitive Scenarios

The study investigates how users perceive the helpfulness and privacy-preservation of large language model (LLM) responses in privacy-sensitive scenarios. Using 94 participants and 90 PrivacyLens scenarios, researchers found that users’ evaluations of identical LLM outputs varied widely, whereas five proxy LLM judges showed high agreement but low correlation with user judgments. The results suggest that proxy LLMs cannot reliably estimate users’ diverse perceptions of utility and privacy, highlighting the need for more user-centered evaluation methods.

By Xiaoyuan Wu, Roshni Kaushik, Wenkai Li, Lujo Bauer, Koichi Onoue
arXiv AI
Aug 7

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

arXiv:2608. 05246v1 Announce Type: new Abstract: Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities.

By Jiahao Zhang, Yongzhi Tong, Zelin Fu, Pengde Zhao, Yanmei Jiang, Jiang Feng, Min Yang
arXiv Computation and Language
Aug 25

PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks

PersonaMem-v3 is a benchmark and evaluation harness designed to assess omni-platform personal intelligence for AI agents. It is built from over one million anonymized real-world engagement histories, covering social media, chatbots, calendars, and AI companions, and tracks user preferences and habits over time. The benchmark tests agents on personalization, LLM-powered recommendation, proactiveness, agentic tool use, and geo-temporal reasoning, evaluating their ability to infer holistic user understanding, personalize responses, rerank recommendations, follow user steering, and avoid inappropriate personalization.

By Bowen Jiang, Yuan Yuan, Zhuoqun Hao, Yuchen Liu, Maohao Shen, Sihao Chen, Gregory Wornell, Chris Callison-Burch, Lyle Ungar, Dan Roth, Qi Guo, Xiangjun Fan, Camillo J. Taylor, Hanchao Yu
arXiv Computation and Language
3d ago

Towards Detecting AI-Assisted Responses in Online Surveys

The paper introduces ASURRE, a benchmark dataset for detecting AI‑assisted responses in online surveys. It evaluates how different LLM usage strategies—ranging from full generation to persona‑grounded agentic completion—affect the performance of existing machine‑generated text detectors. The study finds that while naive AI usage is easily detected, more sophisticated persona‑grounded agents approach chance performance, yet still leave identifiable behavioural traces that can be aggregated to improve detection.

By Qizhou Wang, Bogdan Mamaev, Christopher Leckie
arXiv AI
Aug 11

Autonomy Reshapes How Personalization Affects Privacy Concerns and Trust in LLM Agents

arXiv:2510. 04465v3 Announce Type: replace-cross Abstract: LLM agents require personal information for personalization in order to effectively act on users' behalf, but this raises privacy concerns that can discourage data sharing, limiting both the autonomy levels at which agents can operate and the effectiveness of personalization.

By Zhiping Zhang, Yi Evie Zhang, Freda Shi, Tianshi Li