arXiv AI

Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations

arXiv Machine Learning
1d ago

ZeroHAT: Behavior-Conditioned Zero-Shot Human Activity Trace Generation

ZeroHAT is a new framework for generating synthetic human activity traces (HATs) in a target region without any real data from that region. It transfers behavioral patterns learned from real HATs in source regions and adapts them using publicly available contextual information about the target region. The system includes a consistency-aware intent extractor, a cross-region behavioral cloning module, and a behavior-conditioned activity realization module, and it outperforms the strongest baseline by 4.5–6.4× in downstream utility and improves fidelity by 15.6–40.8% across ten cities.

By Rongchao Xu, Dahai Yu, Lin Jiang, Guang Wang
arXiv AI
1d ago

SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership

SIMLIFE is a scalable platform that simulates long-term household life with rich visual observations, ground-truth action logs, and synthetic dialogues. It introduces the SimLife-BP benchmark, which tests long-context pattern understanding by requiring agents to infer latent behavioral rules from weeks or months of everyday observations across 106 episodes. The benchmark includes 1,439 question-answer pairs that probe direct, counterfactual, noisy, and inverse reasoning under varying rule hints.

By Run Peng, Zinnia Nie, Jing Ding, Yinpei Dai, Yichi Zhang, Zengqing Wu, Yao Fu, Ziqiao Ma, Jiayuan Mao, Joyce Chai
arXiv Computation and Language
Aug 27

EgoArgus: Benchmarking VLMs as Situational Assistants for Modality-Grounded User Supports

EgoArgus is a new, human‑annotated dataset that tests visual‑language models (VLMs) as situational assistants in five everyday dialogue‑video scenarios. It evaluates how well VLMs understand and decide when to intervene, especially when visual and textual cues are helpful, irrelevant, or conflicting. The study finds that current VLMs still struggle to reliably act as egocentric assistants and that existing modality‑bias mitigation methods offer limited improvement.

By Yu-Chien Tang, Yu-Hsiang Liu, An-Zi Yen
arXiv AI
Aug 7

LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs

arXiv:2608. 05246v1 Announce Type: new Abstract: Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities.

By Jiahao Zhang, Yongzhi Tong, Zelin Fu, Pengde Zhao, Yanmei Jiang, Jiang Feng, Min Yang
arXiv AI
Jul 7

Multiplayer Interactive World Models with Representation Autoencoders

arXiv:2607. 05352v1 Announce Type: cross Abstract: We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions.

By Anthony Hu, V\'aclav Volhejn, Adrien Ramanana Rahary, Chris Mulder, Aditya Makkar, Am\'elie Royer, Manu Orsini, Alyx Liao, Adam Jelley, Eloi Alonso, Florian Laurent, Fredrik Nor\'en, James Swingos, Jan H\"unermann, Kent Rollins, Lucas Hosseini, Matthieu Le Cauchois, Maxim Peter, Pim de Witte, Tim Brown, Vincent Micheli, Moritz B\"ohle, Gabriel de Marmiesse, Viktoriia Sharmanska, Lucia Specia, Michael Black, Patrick P\'erez
arXiv AI
Jun 17

Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selective Disclosure

arXiv:2606. 17441v1 Announce Type: cross Abstract: Simulating realistic patient interactions is a key requirement to testing clinical applications of LLMs at scale without time-consuming and expensive user studies.

By Moritz Schlager, Friederike Jungmann, Samuel Schmidgall, Philipp Raffler, Franziska Hartl, Eva Wende, Paula Ro{\ss}m\"uller, Conrad Ketzer, Avinatan Hassidim, Dale R. Webster, Yossi Matias, Yun Liu, Daniel Rueckert, Mike Schaekermann, Paul Hager