arXiv AI

Situation Graph Prediction for User Perspective Modeling

arXiv:2602. 13319v2 Announce Type: replace Abstract: Perspective-aware AI requires modeling evolving internal states---goals, emotions, contexts---not merely preferences.

arXiv Computer Vision
Sep 22

Relationally Grounded Latent World Models for Autonomous Driving

Relationally Grounded Latent World Models for Autonomous Driving proposes using traffic scene graphs as privileged semantic supervision for latent world representations. The approach builds actor‑centric scene graphs from nuScenes 3D annotations, encodes their relational structure with a frozen text embedding model, and aligns visual latent representations to this semantic target during training. At inference the supervision branch is removed, requiring no scene graphs or 3D annotations and adding no extra computation, while achieving a 5.9% reduction in average trajectory L2 error and a 52.4% drop in collision rate compared to the LAW baseline.

By Fabian Schmidt, Markus Enzweiler, Abhinav Valada
arXiv AI
Sep 25

Planning Takes More Than Token Prediction: Causal Plan for Benchmarking and Building Physically Grounded Embodied Reasoners

The paper argues that current embodied vision‑language planning benchmarks favor linguistic next‑token prediction over physically grounded next‑state reasoning, leading models to rely on language priors rather than true causal dependencies. To address this, the authors introduce Causal‑Plan‑Bench, a diagnostic suite covering four causal dimensions, and Causal‑Plan‑1M, a million‑scale corpus of explicit causal reasoning traces extracted from egocentric videos. Extensive experiments show that existing models perform poorly on these tasks, while a new model trained with a tailored recipe—Causal Planner based on Qwen3‑VL‑8B—achieves significant gains, demonstrating the feasibility of physically grounded causal reasoning.

By Zheng Lu, Mingqi Gao, Qinlei Xie, Wanqi Zhong, Hanwen Cui, Zirui Song, Lijie Wang, Chong Luo, Bei Liu, Yiming Li
arXiv AI
Sep 18

SIMLIFE: Pattern Understanding for Long-Horizon Human-Agent Partnership

SIMLIFE is a scalable platform that simulates long-term household life with rich visual observations, ground-truth action logs, and synthetic dialogues. It introduces the SimLife-BP benchmark, which tests long-context pattern understanding by requiring agents to infer latent behavioral rules from weeks or months of everyday observations across 106 episodes. The benchmark includes 1,439 question-answer pairs that probe direct, counterfactual, noisy, and inverse reasoning under varying rule hints.

By Run Peng, Zinnia Nie, Jing Ding, Yinpei Dai, Yichi Zhang, Zengqing Wu, Yao Fu, Ziqiao Ma, Jiayuan Mao, Joyce Chai
arXiv AI
Sep 17

Disentangling Long-Term Memory via Latent Neuro-Symbolic Reasoning

The paper introduces LGM, a neuro‑symbolic framework that disentangles long‑term memory by mapping historical interactions into a continuous latent graph. Instead of static memory graphs, LGM uses a sparse autoencoder to create query‑aware latent nodes and edges, then applies a graph encoder conditioned on the query to perform non‑linear message passing. Experiments on long‑term personalization benchmarks show that LGM outperforms existing methods in capturing both explicit and implicit user preferences and generating personalized responses.

By Cai Ke, Xinghao Chen, Xiaoyu Shen, Keyu Chen, Siyu An, Junnan Dong, Ruifeng Xu, Ruizhi Qiao, Xing Sun
arXiv Computer Vision
6d ago

WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving

WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.

By Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin