arXiv AI By Hiba Eltigani, Rukhshan Haroon, Asli Kocak, Abdullah Bin Faisal, Noah Martin, Fahad Dogar

WaLLM -- Understanding Use and Engagement with a General-Purpose LLM on WhatsApp

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Computation and Language
Aug 25

PersonaMem-v3: Toward Omni-Platform Personal Intelligence for Holistic User Understanding, Recommendation, and Agentic Tasks

PersonaMem-v3 is a benchmark and evaluation harness designed to assess omni-platform personal intelligence for AI agents. It is built from over one million anonymized real-world engagement histories, covering social media, chatbots, calendars, and AI companions, and tracks user preferences and habits over time. The benchmark tests agents on personalization, LLM-powered recommendation, proactiveness, agentic tool use, and geo-temporal reasoning, evaluating their ability to infer holistic user understanding, personalize responses, rerank recommendations, follow user steering, and avoid inappropriate personalization.

By Bowen Jiang, Yuan Yuan, Zhuoqun Hao, Yuchen Liu, Maohao Shen, Sihao Chen, Gregory Wornell, Chris Callison-Burch, Lyle Ungar, Dan Roth, Qi Guo, Xiangjun Fan, Camillo J. Taylor, Hanchao Yu
arXiv AI
Sep 15

Personalizing Personal Health Interfaces: Co-Design with Generative AI

The paper explores how generative AI can lower the barrier to personalizing health dashboards by enabling users to co-design interfaces in Figma Make. In a study with 14 participants, redesigns of Google and Apple Health focused on personal context, future planning, and interactive experiences, though conversational AI designs tended toward chat-window conventions. AI facilitated the materialization of loosely articulated ideas, yet model defaults and generation latency influenced iteration, and the process highlighted interpretability and accountability over privacy, trust, and emotional safety.

By Karthik S. Bhat, Vidhi Shah, Vedika Agnihotri, Dong Whi Yoo, Koustuv Saha