arXiv AI

Harnessing Coupled Stream Completion For Human-Object Interaction Modeling

arXiv Computer Vision
6d ago

RotVLA: Rotational Latent Action for Vision-Language-Action Model

RotVLA introduces a Vision‑Language‑Action framework that replaces discrete latent action encoding with a continuous rotational latent action representation on the group SO(n). This design provides continuity, compositionality, and structured geometry that better capture real‑world action dynamics, and a triplet frame learning scheme enforces meaningful temporal dynamics while preventing degeneration. Trained with 1.7 B parameters on large cross‑embodiment datasets, RotVLA achieves state‑of‑the‑art performance on LIBERO and RoboTwin2.0 benchmarks and shows strong real‑world manipulation results.

By Qiwei Li, Xicheng Gong, Xinghang Li, Peiyan Li, Quanyun Zhou, Hangjun Ye, Jiahuan Zhou, Yadong Mu
arXiv Computer Vision
Sep 21

GALA: Geometry-Aware Latent Action Modeling for Vision-Language-Action Model Pretraining across Embodiments

The paper introduces GALA, a Geometry-Aware Latent Action modeling framework that enhances image-based latent actions with 3D end‑effector motion. It proposes the Unified End‑effector Motion Representation (UEMR) to preserve fine‑grained motion while improving cross‑embodiment generalizability. Experiments show GALA effectively models generalizable fine‑grained motions across embodiments, achieving high success rates in RoboCasa-GR1 and real‑world tasks.

By Yichen Liu, Puzhen Yuan, Xiang Zhu, Yanjiang Guo, Jianyu Chen
arXiv AI
1d ago

Generative Interactions: Weaving Multiparty Human Motion with Bilevel Latent Dynamics

The paper introduces BRAID, a hierarchical latent-variable model that generates multi-person human motion by explicitly modeling both group-level interaction dynamics and individual behavior conditioned on evolving group context. It treats social motion generation as a meta-transfer learning problem, learning shared interaction priors across datasets and adapting them to arbitrary context sets of observed people and joints. BRAID supports coherent generation under full, sparse, or partial observations and produces compact social-state vectors useful for downstream embodied-agent systems, with evaluations on social forecasting, tracking, in-filling, and response generation.

By Ojas Shirekar, Yash Surange, Agustinas Ju\v{c}as, Chirag Raman
arXiv Computer Vision
Aug 28

Residual Flow Matching with Dynamic Cross-Interaction for 3D Multi-Person Motion Prediction

The paper introduces a Prior‑Guided Residual Flow Matching framework for 3D multi‑person motion prediction. It uses a Deterministic Coarse Prior to anchor kinematics and a Dynamic Cross‑Interaction mechanism to synchronize inter‑agent message passing during integration, thereby improving structural consistency and social context extraction. A decoupled joint‑motion architecture with bidirectional fusion further preserves fine‑grained kinematic coherence, achieving state‑of‑the‑art accuracy on several datasets.

By Wei Wei, Yinyuan Zhao, Ruixuan Yu
arXiv Computer Vision
Aug 27

InteractGesture: Progressive Chunk Guidance for Continuous Streaming Co-Speech Gesture Control

InteractGesture is a model‑agnostic, inference‑time method that enables fine‑grained spatial control of individual joints in continuous streaming co‑speech gesture generation. It guides diffusion sampler latent estimates through a differentiable RVQ‑VAE decoder, backpropagating spatial control gradients to adjust motion latents during sampling. To address chunk‑wise dependency issues in streaming generation, the method introduces Progressive Chunk Guidance, a chunk‑window strategy that keeps an active set of editable chunk latents with staggered delays, allowing spatial constraints to propagate gradients backward across chunk boundaries and reducing boundary inconsistencies.

By Ekkasit Pinyoanuntapong, Ajinkya Deogade, Paul Streli, Wenjing Zhang, Joanna Materzynska, Pu Wang, Vittorio Ferrari, Jie Shen
arXiv Computer Vision
Sep 15

PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models

PhysBrain 1.5 is a unified vision‑language model that learns to understand physical environments, generate actions, and predict future states by encoding language, end‑effector motion, and dense visual targets as discrete sequences and training them with autoregressive next‑token prediction. The model is pre‑trained on human interaction videos and fine‑tuned on human demonstrations, robot trajectories, and simulated experience, achieving an average score of 72.5 across 28 embodied understanding benchmarks and outperforming other open‑source models on 14 of them. It also demonstrates the ability to produce end‑effector trajectories and predict future scenes with spatially aligned RGB, depth, and robot‑mask outputs.

By DeepCybo Team, Yu Bin, Haipeng Cao, Zheng Chang, Kai Chen, Youning Chen, Kailin Deng, Yichao Du, Xiaotong Fu, Haoyang Ge, Yunlong Guo, Chenliu Hao, Jiyan He, Xuguo He, Yakun Hou, Kai Hu, Cong Huang, Tuopusen Huang, Yu Huang, Hong Li, Peize Li, Shijie Lian, Xiaopeng Lin, Yun Lin, Haibao Liu, Haochen Liu, Qiuzhi Liu, Shengcai Liu, Zhiqiang Liu, Tao Luo, Peng Ren, Shuo Ren, Chaoyi Ruan, Zhaolong Shen, Yukun Shi, Qiyuan Su, Yuxuan Tian, Yining Wang, Changti Wu, Hao Wu, Xueyin Xu, Ruoqi Yang, Zhaoyang Yang, Hang Yuan, Zhaoyang Zeng, Hanwen Zhang, Ruimeng Zhang, Yao Zhang, Yibo Zhang, Yuxiang Zhang, Zhirui Zhang, Ziyi Zhang, Zubin Zheng, Zishen Zhuang