arXiv AI By Run Peng, Zinnia Nie, Jing Ding, Yinpei Dai, Yichi Zhang, Zengqing Wu, Yao Fu, Ziqiao Ma, Jiayuan Mao, Joyce Chai

SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership

Read the original on arXiv AI →

SIMLIFE is a scalable platform that simulates long-term household life with rich visual observations, ground-truth action logs, and synthetic dialogues. It introduces the SimLife-BP benchmark, which tests long-context pattern understanding by requiring agents to infer latent behavioral rules from weeks or months of everyday observations across 106 episodes. The benchmark includes 1,439 question-answer pairs that probe direct, counterfactual, noisy, and inverse reasoning under varying rule hints.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 2

HugAgent: A Human Simulation Benchmark for Individual-Level Reasoning

arXiv:2510.15144v4 Announce Type: replace Abstract: Simulating human reasoning in open-ended tasks has long been a central aspiration in AI and cognitive science. While large language models now appr...

By Chance Jiajie Li, Zhenze Mo, Yuhan Tang, Ao Qu, Jiayi Wu, Kaiya Ivy Zhao, Yulu Gan, Jie Fan, Jiangbo Yu, Hang Jiang, Paul Pu Liang, Jinhua Zhao, Luis Alberto Alonso Pastor, Kent Larson
arXiv AI
Sep 4

Imagine-then-Plan: Agent Learning from Adaptive Lookahead with World Models

The paper introduces Imagine-then-Plan (ITP), a framework that lets agents learn by interacting with a learned world model to generate multi-step imagined trajectories. ITP features an adaptive lookahead mechanism that balances ultimate goals with task progress, producing richer signals about future outcomes. Experiments on various benchmarks show that ITP outperforms existing baselines, and analyses suggest the adaptive lookahead improves reasoning for complex tasks.

By Youwei Liu, Jian Wang, Hanlin Wang, Beichen Guo, Wenjie Li