arXiv AI

SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data

arXiv:2607. 19949v1 Announce Type: new Abstract: Smartphone personal assistants reason over longitudinal personal data, yet evaluating them requires context-rich evaluation data whose correct answers are known, and real device traces are too privacy-sensitive to share.

arXiv AI
Aug 12

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

arXiv:2608. 10692v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions.

By Junjie Ye, Zhuohui Sheng, Shaofan Liu, Yulun Zhu, Wenjie Fu, Dingwei Zhu, Ming Zhang, Yujiong Shen, Weichao Wang, Xin Zhao, Shihan Dou, Tao Gui, Qi Zhang, Xuanjing Huang, Pluto Zhou
arXiv AI
Aug 25

CallScreenBench: Benchmarking Small Language Models as Phone Secretaries

CallScreenBench is a benchmark for evaluating small, on-device language models that act as phone secretaries, focusing on their ability to handle unknown inbound calls without a cooperative task. The benchmark measures owner endorsement through five call-and-note metrics, each paired with counter-metrics and uncertainty estimates, and includes guardedness diagnostics to identify safe, tool‑free proxies. Results across 4‑bit checkpoints of 0.6‑4 B parameter models show varying performance on service, recall, plausibility, and triage discrimination, highlighting trade‑offs between quality and guardedness.

By Jiaqi Gan, Haoyuan Tang, Jamey Z. Liang, Siying Chen, Ankit Raj, Kidus Zewde, Yuchen Zhou, Yuxin Zhang, Simiao Ren
arXiv AI
Aug 17

PhoneWorld: Scaling Phone-Use Agent Environments

arXiv:2605. 29486v2 Announce Type: replace-cross Abstract: A central bottleneck for phone-use agents is that controllable, reproducible environments covering real mobile behavior are hard to build at scale.

By Yuxuan Liu, Xin Lai, Junyi Li, Pengyuan Lyu, Jason, Yiduo Guo, Zhengyao Fang, Yang Ding, Yi Zhang, Weinong Wang, Huawen Shen, Xingran Zhou, Liang Wu, Fei Tang, Sunqi Fan, Shangpin Peng, Zheng Ruan, Anran Zhang, Chengquan Zhang, Han Hu, Benyou Wang, Ji-Rong Wen, Rui Yan, Zhengyang Tang
arXiv AI
Sep 10

A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation

The paper introduces a three-tier persona vector for user simulation in evaluating LLM agents, comprising 23 dimensions across demographics, behavioral traits, and emotional states, plus a query-complexity overlay. It demonstrates that these nuanced personas generate diverse, scenario-reactive conversations, leading to significant variations in agent goal achievement and compliance across different contexts. The model’s design allows for reproducible, auditable user behavior patterns without relying on learned covariance matrices.

By Rahul Khedar, Eshita, Sneha Teja Sree Reddy Thondapu, Mayank Malhotra, Arup Kumar Das, Jitesh Chandra Mishra, Arun Menon, Avinash Karn, Mouli V
arXiv AI
Sep 18

MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

MTVA-Bench is a new benchmark designed to evaluate the language model component of cascaded voice agents under realistic conditions. It simulates callers with an LLM, mocks backend tool responses, and scores both tool‑call correctness and conversational quality using two LLM judges, covering 49 agents, 490 scenarios, and 7 languages. The benchmark reveals that while models perform similarly on tool selection, they differ widely in argument handling, rule compliance, and dialogue quality, highlighting the nuanced challenges of real‑world voice interactions.

By Pritish Mishra, Ishaan Kumar, Akshat Mandoli, Sudarshan Kamath
arXiv Computation and Language
Aug 28

ContextEcho: A Benchmark for Persona Drift in Long Agentic-Coding Sessions

ContextEcho is a benchmark and harness designed to measure persona drift in large language models during long, tool‑using coding sessions. It includes a 25‑probe identity suite, a snapshot‑then‑probe protocol that preserves the main conversation, and both judged and judge‑free measurement surfaces. Across 23 frontier models and thousands of turns, the benchmark shows that persona drift is widespread, not limited to specific model families, and that simple in‑session compaction does not reset it, while a single‑shot anchor can restore the intended persona.

By Xianzhong Ding, Yangyang Yu, Changwei Liu, Bill Zhao, Le Chen, Tao Chen
Hugging Face Trending Papers
Sep 8

A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation

The paper introduces a three-tier persona vector to generate diverse, realistic user inputs for evaluating tool-augmented LLM agents. The vector includes 23 dimensions: categorical demographics, continuous behavioral traits, and continuous emotional states, plus a query-complexity overlay. Experiments on 64,698 conversations show that these persona dimensions produce measurable differences in agent performance and realistic scenario-reactive behavior.

arXiv AI
Sep 24

MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research

MobileGym is a browser-hosted, lightweight simulation platform designed for mobile GUI agent research. It offers verifiable outcome signals via deterministic, JSON-based state judging and supports scalable online reinforcement learning with hundreds of parallel instances on a single server. The platform includes a declarative task-definition framework, a structured AnswerSheet protocol, and a benchmark of 416 parameterized tasks across 28 apps, demonstrating strong sim-to-real transfer in a case study.

By Dingbang Wu, Rui Hao, Haiyang Wang, Shuzhe Wu, Han Xiao, Zhenghong Li, Bojiang Zhou, Zheng Ju, Zichen Liu, Lue Fan, Zhaoxiang Zhang