arXiv:2608.24334v1 Announce Type: new
Abstract: Discrete motion representations have substantially advanced autoregressive text-to-motion generation. However, most motion tokenizers are optimized for...
By Tianlv Huang, Hetian Guo, Ziyi Cai, Song Wang, Yanping Zhang, Zipei Fan, Xuan Song, Guangming Wu, Xin Zheng
arXiv:2607. 08741v1 Announce Type: cross Abstract: Generating realistic 3D human motions in real-time within interactive applications is key for animation, simulation, and humanoid robotics.
By Kaifeng Zhao, Mathis Petrovich, Haotian Zhang, Tingwu Wang, Siyu Tang, Davis Rempe
The paper introduces LaPla, a Vision‑Language‑Action framework that uses a latent‑aligned planning approach to convert discrete semantic reasoning into continuous, physics‑constrained driving actions. It employs a residual VQ‑VAE to encode vehicle kinematics into a structured latent space, then projects multimodal inputs—images, past actions, and text—directly into this latent space, allowing a frozen decoder to generate physically plausible trajectories without quantization errors. Experiments on nuScenes and NVIDIA AlpaSim show LaPla reduces long‑horizon L2 error by 15.52% and improves closed‑loop success rates by 33.34 percentage points while cutting inference latency.
By Ruoyu Yao, Yusen Xie, Qingzhao Liu, Pei Liu, Zewei Yang, Yipeng Zhu, Xiaolong Wang, Jun Ma
arXiv:2609.14615v1 Announce Type: cross
Abstract: Unified motion generation and understanding is crucial for embodied AI systems that can both synthesize and interpret human actions in open-world env...
By Guocun Wang, Kenkun Liu, Guorui Song, Jing Lin, Zhe Huang, Luyuan Zhang, Dake Zhong, Choo Sin Wai, Xiaoguang Han, Haoqian Wang
arXiv:2606. 30266v1 Announce Type: cross Abstract: Motion-language agents must possess the bidirectional capability to both understand human movement (motion-to-text, M2T) and generate it from natural language (text-to-motion, T2M).
By Bertram Taetz, Hugo Albuquerque Cosme da Silva, Gabriele Bleser-Taetz
arXiv:2608.23279v1 Announce Type: new
Abstract: Text-driven human motion synthesis has made substantial development with two core modules of motion representation and generative architecture. For rep...
By Chengqun Yang, Liang Xu, Yanping Li, Fulong Liu, Jingnan Gao, Weili Zeng, Yichao Yan
RotVLA introduces a Vision‑Language‑Action framework that replaces discrete latent action encoding with a continuous rotational latent action representation on the group SO(n). This design provides continuity, compositionality, and structured geometry that better capture real‑world action dynamics, and a triplet frame learning scheme enforces meaningful temporal dynamics while preventing degeneration. Trained with 1.7 B parameters on large cross‑embodiment datasets, RotVLA achieves state‑of‑the‑art performance on LIBERO and RoboTwin2.0 benchmarks and shows strong real‑world manipulation results.
By Qiwei Li, Xicheng Gong, Xinghang Li, Peiyan Li, Quanyun Zhou, Hangjun Ye, Jiahuan Zhou, Yadong Mu
arXiv:2603.08590v4 Announce Type: replace
Abstract: Text-to-motion generation has advanced with larger corpora and stronger generators, yet many models still rely on holistic frame- or clip-level lat...
By Zeyu Ling, Qing Shuai, Teng Zhang, Shiyang Li, Bo Han, Changqing Zou
ReMoMask-2 is a retrieval‑augmented text‑to‑motion generation framework that improves on complex motion descriptions by addressing coarse retrieval and representation gaps. It introduces a structure‑aware RAG pipeline with Hierarchical Bidirectional Momentum contrastive learning, Semantic Spatial‑Temporal Attention, and Topology Structured Masking, and rebuilds the retrieval database in the generator’s latent space using a lightweight projector. Experiments on HumanML3D, KIT‑ML, and SnapMoGen show state‑of‑the‑art retrieval accuracy and the lowest FID scores, with a single mask‑transformer stage delivering faster inference than the previous two‑stage design.
arXiv:2608.18734v2 Announce Type: replace
Abstract: 4D understanding and reasoning is a fundamental capability for embodied AI agents operating in dynamic physical environments. However, existing vis...
By Kumal Hewagamage, Isuranga Senavirathne, Sasika Amarasinghe, Hasitha Gallella, Dulanga Weerakoon, Vigneshwaran Subbaraju, Ranga Rodrigo
arXiv:2603. 22282v2 Announce Type: replace-cross Abstract: We present UniMotion, to our knowledge the first unified framework for simultaneous understanding and generation of human motion, natural language, and RGB images within a single architecture.
By Ziyi Wang, Xinshun Wang, Shuang Chen, Yang Cong, Mengyuan Liu
ZYT-World is a real‑time, controllable world model designed for closed‑loop autonomous‑driving simulation. It generates a full fisheye‑pinhole camera rig at native resolution, using Plucker adapters for camera geometry, ego‑motion adaptive layer normalization for motion control, and a lightweight pixel‑aligned layout to condition traffic participants. The model achieves high fidelity to a 40‑step teacher while being over 100 times faster, and includes an implicit‑memory module that preserves place‑specific evidence for long‑horizon stability.
By Boni Hu, Xiong Wei, Haoming Huang, Yong Huang, Chenbo Wang, Yi Yang, Jiancheng Wang, Ruicheng Zhu, Zhimin Yang, Guanglai Liu, Qiaowan Jin, Dongzhuo Wang, Haiwei Kuang, Jiajun Fan, Yue Wu, Jiaxin Wei, Hao Sun, Feihong Yan, Wei Bi, Kaixuan Wang, Zichao Guo, Xiaozhi Chen