arXiv:2608. 06657v1 Announce Type: new Abstract: Modern cyber-physical and AI-assisted systems couple human operators, AI decision modules, and automated controllers in a single control loop, so trustworthiness depends on the whole loop, not any one model.
By Joshua Zuniga, Srinivasan Subramanian, Ramya Madhuri Narapureddy, Md Abdullah Al Hafiz Khan
MPCFormer is a physics‑informed, data‑driven framework that explicitly models multi‑vehicle social interaction dynamics for autonomous driving. It uses a Transformer‑based encoder‑decoder to learn discrete state‑space dynamics from naturalistic data, enabling explainable, human‑like behavior planning within a Model Predictive Control (MPC) framework. In open‑loop NGSIM tests, it achieves the lowest trajectory prediction errors (ADE 0.86 m over 5 s), and in closed‑loop intense interaction scenarios it attains a 94.67 % planning success rate, 15.75 % efficiency gain, and reduces collisions from 21.25 % to 0.5 %.
By Jia Hu, Zhexi Lian, Xuerun Yan, Ruiang Bi, Dou Shen, Yu Ruan, Chunlong Xia, Haoran Wang
arXiv:2607. 27017v1 Announce Type: new Abstract: A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment.
By Kaizhen Tan (New York University, Carnegie Mellon University), Xin Xu (Carnegie Mellon University), Siru Tao (Carnegie Mellon University), Hanzhe Hong (Carnegie Mellon University), Yang Feng (Columbia University), Heqing Du (Columbia University)
Traffic microsimulators rely on hand-crafted behavior models that reproduce aggregate flow but miss the heterogeneous interactions between vehicles at signalized intersections. Learned trajectory predictors capture richer interactions but are short-horizon and tend to be unstable when run in closed loop.
Kairos extends a hierarchical 3D scene graph to a 4D scene graph, storing for each voxel a directional mixture and a presence rate. Spectral predictors forecast both the probability of people being present and the full directional distribution of their motion at any future query time. The model supports conditional queries via pairwise flow dependence and provides calibrated credible intervals that tighten as observations accumulate.
By Iacopo Catalano, Julio A. Placed, Javier Civera, Jorge Pe\~na Queralta
Long-term autonomy in human-populated environments requires anticipating whether and how people will move at times a robot has not yet observed. Existing representations of pedestrian motion face a tr...
The paper proposes a method to distill world‑model representations into compact Vision‑Language‑Action (VLA) policies. By adding a single feature‑alignment term during VLA training, a frozen world model’s internal features are cached and the student policy learns to match them, eliminating the need for a generative future‑rolling component. The resulting lightweight policy runs in 32 ms on an RTX 5090, achieving high performance on LIBERO and RoboCasa‑GR1, and transfers effectively to real robotic hardware.
By Trung Dao, Sankalp Yamsani, Jaden Park, Joohyung Kim, Yong Jae Lee
PV-WM is a history‑only world model that jointly predicts pedestrian root motion, 15‑joint articulation, and vehicle kinematic states in a synchronized heterogeneous state. It uses recurrent updates to generate pedestrian and vehicle motion chunks, reconstructing vehicle boxes from predicted center, heading, and observed extent, and recomputes pedestrian‑vehicle geometry after each transition. Compared to a one‑shot predictor, PV‑WM reduces Root ADE by 12.7% and MPJPE by 14.8%, and across 824 Waymo contexts it lowers Root ADE by 5.2%, MPJPE by 7.6%, P‑V distance error by 11.9%, and oriented‑box closest‑approach error by 5.8%, while using 57.1% fewer parameters, 96.5% fewer FLOPs, and 25.5% lower p95 latency.
By Haozhuang Chi, Jingsong Liang, Ziying Song, Lei Yang, Shihao Li, Haoruo Zhang, Chen Lv
GigaBrain-WBC-0.5 is a Behavior World Model that uses a causal Transformer to predict next actions, states, and a distribution over latent behavior commands for humanoid whole-body control. It incorporates an automatic terrain-annotation pipeline to recover 3D contact geometry from motion data, allowing the model to learn how terrain and objects influence dynamics. The system detects implausible commands online, retracts them onto learned behaviors, and achieves high success rates in terrain interaction, command robustness, and fall recovery, with promising hardware trials on different robots.
By Ziyang Cheng, Tianshu Tang, Jinxin Lan, Xinze Chen, Yuhan Gong, Zhichao Liu, Changzhong Wu, Yahao Mao, Zongyan Deng, Mingxuan Ma, Huasen Xi, Yilong Liu, Yutong Wu, Xiaofeng Wang, Yang Wang, Yun Ye, Guan Huang, Xiaojie Jin, Zheng Zhu, Jiwen Lu
arXiv:2606. 16605v1 Announce Type: new Abstract: World models are widely used in robotic and agentic engineering control systems due to their ability to learn latent dynamics for planning and decision-making.
By Junjian Zhang, Hao Tan, Ruonan Li, Dong Zhu, Aiping Li, Zhaoquan Gu
arXiv:2606. 00090v1 Announce Type: cross Abstract: Physical AI systems increasingly map multimodal observations, language instructions, and learned world representations into physically consequential actions.
By Barak Or
arXiv:2608. 01587v1 Announce Type: cross Abstract: Machine-learning benchmarks often pair a label that aggregates a long temporal horizon with input observed through one or a few short windows.
By Xizhe Zhang