We present FlowHMR, a framework for recovering physically plausible global 3D human motion from monocular video. Previous learning-based methods typically regress human motion directly from video and...
World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for control raises a new challenge: how to represent actions in a suitable form that aligns with pretrained video generators while carrying enough motion cues for accurate control.
LPA-CWM introduces a Learned Physical Adjudicator (LPA) to improve counterfactual world models (CWM) for motion reasoning by learning to weight candidate responses based on visual context and response structure. The 3.0M‑parameter LPA is trained on dense MOVi‑F trajectories while keeping the CWM predictor and intervention generator frozen. A new Completeness‑aware Motion Correspondence (CMC) protocol evaluates localization, trajectory completeness, visibility, and continuity, and LPA‑CWM achieves significant gains on DAVIS and Kinetics subsets.
By Kunwei Wu, Xiang Liu, Guocai Yao, Junming Chen, Zhikang Chen, Min Zhang, Pengwei Wang, Sen Cui
arXiv:2609.17521v1 Announce Type: cross
Abstract: Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipulation of dynamic scenes. Yet...
By Chuhao Chen, Peter Wonka, Chaoyang Wang, Chen Wang, Qiao Feng, Sergey Tulyakov, Lingjie Liu
arXiv:2608. 19556v1 Announce Type: cross Abstract: Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a coherent world: long rollouts accumulate geometric drift and degrade into static or unnatural motion.
By Yuanhao Ban, Jiaqi Feng, Hengguang Zhou, Xiaohuan Pei, Justin Cui, Cho-Jui Hsieh
Video Prediction Policy 2 (VPP2) is a new world action model that improves zero‑shot generalization for both video prediction and action generation. It achieves this by pretraining a large, diverse manipulation video dataset with event‑level supervision, then distilling the model into a single‑step visual planner and adding a mixture‑of‑transformers action module. Experiments show VPP2 outperforms leading baselines on open‑ended video prediction, real‑world zero‑shot manipulation, and several challenging benchmarks.
By Yanjiang Guo, Haodong Yan, Zhide Zhong, Zhongru Zhang, Qingyuan Yang, Qingzhou Lu, Xiaoyu Chen, Yen-Jen Wang, Shuying Deng, Chenghan Yang, Puzhen Yuan, Chenxin Liu, Tun Ban, Xiang Zhu, Yichen Liu, Kun Feng, Haoang Li, Jianyu Chen
The paper introduces a reinforcement learning post‑training scheme that trains robot world models on their own autoregressive rollouts, using a contrastive RL objective adapted from diffusion models. It also proposes a training protocol that compares multiple variable‑length futures, a multi‑view visual fidelity reward, and demonstrates state‑of‑the‑art rollout fidelity on the DROID dataset, outperforming baselines on LPIPS, SSIM, and human preference tests.
By Jai Bardhan, Patrik Drozdik, Josef Sivic, Vladimir Petrik
The paper introduces 4DStreamCtrl, a system that unifies camera motion, object trajectories, and depth into a single 3D point‑track representation, enabling joint control, depth editing, and motion transfer in a single forward pass. By mining in‑the‑wild video for 3D motion supervision and encoding it with a lightweight Geometric Motion Head, the authors train a causal streaming student that can generate arbitrarily long videos in just four denoising steps, achieving 20 FPS on a single high‑end GPU for 480p video. This approach outperforms prior camera‑only, 2D, and offline‑3D methods in motion‑control precision while maintaining temporal coherence over hundreds of frames, thereby enabling interactive 4D‑controllable streaming generation for the first time.
By Shiqian Li, Chenguo Lin, Zhiguang Liu, Yu Tang, Jiarong Ou, Rui Chen, Yixin Zhu
arXiv:2609.19142v1 Announce Type: new
Abstract: World models endow perceptual systems with the ability to predict how scenes evolve under interaction. They are most beneficial when trained on diverse...
By Bardienus P. Duisterhof, Kaifeng Zhang, Adam Hung, Bowen Wen, Stan Birchfield, Yunzhu Li, Deva Ramanan, Jeffrey Ichnowski
Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a coherent world: long rollouts accumulate geometric drift and degrade into static or unnatural motion. Recent bidirectional approaches address this problem using rewards signals built upon 3D Gaussian-Splatting reconstruction.
Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipulation of dynamic scenes. Yet existing controllable methods either require the f...
PhysWAM is a unified world-action model for autonomous driving that jointly denoises multiview video, metric depth, and ego motion using a flow‑matching transformer. It introduces Coupled Point Projection (CPP), a geometric constraint that aligns generated depth points with LiDAR data after applying the predicted SE(3) ego motion, thereby enforcing physical consistency. At inference, trajectory selection uses a simple label‑free consensus rule, and the model demonstrates strong planning performance, zero‑shot transfer to unseen environments, and accurate, temporally coherent depth and video predictions.
By Dhruv Parikh, Fengcheng Yu, Quankai Gao, Jiawei Yang, Junjie Ye, Maulik Bhatt, Thang Vu, Charles Ochoa, Rowan McAllister, Igor Vasiljevic, Rajgopal Kannan, Viktor Prasanna, Vitor Guizilini, Yue Wang