Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2607. 29181v1 Announce Type: cross Abstract: Agentic assistants capable of proactive, personalized interactions require structured models of user intent and workflow.
ZeroHAT is a new framework for generating synthetic human activity traces (HATs) in a target region without any real data from that region. It transfers behavioral patterns learned from real HATs in source regions and adapts them using publicly available contextual information about the target region. The system includes a consistency-aware intent extractor, a cross-region behavioral cloning module, and a behavior-conditioned activity realization module, and it outperforms the strongest baseline by 4.5–6.4× in downstream utility and improves fidelity by 15.6–40.8% across ten cities.
SIMLIFE is a scalable platform that simulates long-term household life with rich visual observations, ground-truth action logs, and synthetic dialogues. It introduces the SimLife-BP benchmark, which tests long-context pattern understanding by requiring agents to infer latent behavioral rules from weeks or months of everyday observations across 106 episodes. The benchmark includes 1,439 question-answer pairs that probe direct, counterfactual, noisy, and inverse reasoning under varying rule hints.
EgoArgus is a new, human‑annotated dataset that tests visual‑language models (VLMs) as situational assistants in five everyday dialogue‑video scenarios. It evaluates how well VLMs understand and decide when to intervene, especially when visual and textual cues are helpful, irrelevant, or conflicting. The study finds that current VLMs still struggle to reliably act as egocentric assistants and that existing modality‑bias mitigation methods offer limited improvement.
arXiv:2608. 05246v1 Announce Type: new Abstract: Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities.
arXiv:2606. 12657v1 Announce Type: new Abstract: Human mobility data is important for transportation, urban planning, and epidemic control, but large-scale trajectory collection is often costly and privacy-constrained, motivating realistic synthetic trajectory generation.