arXiv AI

BEHAVE: Real-Time Modeling of Human Systems as Observable Complex Dynamical Systems and Operational Objects for Physical AI

BEHAVE models an interacting human group as a complex dynamical system called a HumanSystem, whose state is partly encoded in the interaction structure rather than individual tracks. By incorporating interaction evidence, the method improves group discrimination and captures differences in neighbor-level organization during bottleneck scenarios. The framework derives routing, local dynamics, and stability metrics, enabling real-time querying of group state and critical modes for Physical AI applications.

arXiv AI
Aug 25

MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving

MPCFormer is a physics‑informed, data‑driven framework that explicitly models multi‑vehicle social interaction dynamics for autonomous driving. It uses a Transformer‑based encoder‑decoder to learn discrete state‑space dynamics from naturalistic data, enabling explainable, human‑like behavior planning within a Model Predictive Control (MPC) framework. In open‑loop NGSIM tests, it achieves the lowest trajectory prediction errors (ADE 0.86 m over 5 s), and in closed‑loop intense interaction scenarios it attains a 94.67 % planning success rate, 15.75 % efficiency gain, and reduces collisions from 21.25 % to 0.5 %.

By Jia Hu, Zhexi Lian, Xuerun Yan, Ruiang Bi, Dou Shen, Yu Ruan, Chunlong Xia, Haoran Wang
arXiv Machine Learning
Jul 30

What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

arXiv:2607. 27017v1 Announce Type: new Abstract: A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment.

By Kaizhen Tan (New York University, Carnegie Mellon University), Xin Xu (Carnegie Mellon University), Siru Tao (Carnegie Mellon University), Hanzhe Hong (Carnegie Mellon University), Yang Feng (Columbia University), Heqing Du (Columbia University)
arXiv AI
Sep 24

Kairos: Grounded Forecasting of Presence and Directional Flow in 4D Scene Graphs

Kairos extends a hierarchical 3D scene graph to a 4D scene graph, storing for each voxel a directional mixture and a presence rate. Spectral predictors forecast both the probability of people being present and the full directional distribution of their motion at any future query time. The model supports conditional queries via pairwise flow dependence and provides calibrated credible intervals that tighten as observations accumulate.

By Iacopo Catalano, Julio A. Placed, Javier Civera, Jorge Pe\~na Queralta
arXiv Computer Vision
Sep 22

Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies

The paper proposes a method to distill world‑model representations into compact Vision‑Language‑Action (VLA) policies. By adding a single feature‑alignment term during VLA training, a frozen world model’s internal features are cached and the student policy learns to match them, eliminating the need for a generative future‑rolling component. The resulting lightweight policy runs in 32 ms on an RTX 5090, achieving high performance on LIBERO and RoboCasa‑GR1, and transfers effectively to real robotic hardware.

By Trung Dao, Sankalp Yamsani, Jaden Park, Joohyung Kim, Yong Jae Lee
arXiv AI
Sep 10

PV-WM: A Heterogeneous Micro-Macro World Model for Articulated Pedestrian-Vehicle Co-Rollout

PV-WM is a history‑only world model that jointly predicts pedestrian root motion, 15‑joint articulation, and vehicle kinematic states in a synchronized heterogeneous state. It uses recurrent updates to generate pedestrian and vehicle motion chunks, reconstructing vehicle boxes from predicted center, heading, and observed extent, and recomputes pedestrian‑vehicle geometry after each transition. Compared to a one‑shot predictor, PV‑WM reduces Root ADE by 12.7% and MPJPE by 14.8%, and across 824 Waymo contexts it lowers Root ADE by 5.2%, MPJPE by 7.6%, P‑V distance error by 11.9%, and oriented‑box closest‑approach error by 5.8%, while using 57.1% fewer parameters, 96.5% fewer FLOPs, and 25.5% lower p95 latency.

By Haozhuang Chi, Jingsong Liang, Ziying Song, Lei Yang, Shihao Li, Haoruo Zhang, Chen Lv
arXiv AI
Aug 20

GigaBrain-WBC-0.5: A Behavior World Model for Robust Whole-Body Control with Environment Interaction

GigaBrain-WBC-0.5 is a Behavior World Model that uses a causal Transformer to predict next actions, states, and a distribution over latent behavior commands for humanoid whole-body control. It incorporates an automatic terrain-annotation pipeline to recover 3D contact geometry from motion data, allowing the model to learn how terrain and objects influence dynamics. The system detects implausible commands online, retracts them onto learned behaviors, and achieves high success rates in terrain interaction, command robustness, and fall recovery, with promising hardware trials on different robots.

By Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Yang Wang, Yun Ye, Guan Huang, Xiaojie Jin, Zheng Zhu, Jiwen Lu