FlowHMR: Physically Plausible Motion Capture from Video
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
We present FlowHMR, a framework for recovering physically plausible global 3D human motion from monocular video. Previous learning-based methods typically regress human motion directly from video and...
World Action Models (WAMs) are able to leverage pretrained video generators for both world modeling and action prediction. However, directly leveraging such video generators for control raises a new challenge: how to represent actions in a suitable form that aligns with pretrained video generators while carrying enough motion cues for accurate control.
LPA-CWM introduces a Learned Physical Adjudicator (LPA) to improve counterfactual world models (CWM) for motion reasoning by learning to weight candidate responses based on visual context and response structure. The 3.0M‑parameter LPA is trained on dense MOVi‑F trajectories while keeping the CWM predictor and intervention generator frozen. A new Completeness‑aware Motion Correspondence (CMC) protocol evaluates localization, trajectory completeness, visibility, and continuity, and LPA‑CWM achieves significant gains on DAVIS and Kinetics subsets.
arXiv:2609.17521v1 Announce Type: cross Abstract: Interactive control for video generation is moving from coarse prompts toward fine-grained, physically meaningful manipulation of dynamic scenes. Yet...
arXiv:2608. 19556v1 Announce Type: cross Abstract: Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a coherent world: long rollouts accumulate geometric drift and degrade into static or unnatural motion.
Video Prediction Policy 2 (VPP2) is a new world action model that improves zero‑shot generalization for both video prediction and action generation. It achieves this by pretraining a large, diverse manipulation video dataset with event‑level supervision, then distilling the model into a single‑step visual planner and adding a mixture‑of‑transformers action module. Experiments show VPP2 outperforms leading baselines on open‑ended video prediction, real‑world zero‑shot manipulation, and several challenging benchmarks.