AdvMT: Adversarial Motion Transformer for Long-term Human Motion Prediction
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2607. 08725v1 Announce Type: cross Abstract: Recent progress in 3D human pose estimation has made markerless recovery of skeletal motion increasingly accurate and scalable.
Pose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming proposes PoseOFF, a representation that captures local motion around human joints by conditioning optical flow extraction on pose. This structured motion representation aligns with human kinematics and improves early action recognition accuracy across multiple datasets and backbones. PoseOFF achieves comparable or better performance while observing less of the action sequence, making it suitable for real‑time, resource‑constrained robotic systems.
The key challenge in articulated 3D object generation from a single image is accurately predicting the underlying kinematic structure. Existing methods either infer kinematic parameters directly from a static image that lacks dynamic part-level kinematic relationships, or estimate parameters from visual dynamics generated from a single image, which is prone to accumulated errors of two steps.
arXiv:2607. 08741v1 Announce Type: cross Abstract: Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics.
arXiv:2608. 16222v1 Announce Type: cross Abstract: Humanoid intelligence requires learning over an extremely diverse space of whole-body motions and physically grounded interactions.
The paper introduces a Prior‑Guided Residual Flow Matching framework for 3D multi‑person motion prediction. It uses a Deterministic Coarse Prior to anchor kinematics and a Dynamic Cross‑Interaction mechanism to synchronize inter‑agent message passing during integration, thereby improving structural consistency and social context extraction. A decoupled joint‑motion architecture with bidirectional fusion further preserves fine‑grained kinematic coherence, achieving state‑of‑the‑art accuracy on several datasets.