arXiv AI

Towards Human Motion World Models via Executable Behaviour Representations

arXiv:2604. 18064v2 Announce Type: replace Abstract: Human motion world models should capture motion's intentionality by being executable: adaptable to different actions and capable of assessing motion quality.

arXiv Computer Vision
Aug 27

Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

Zero-WAM introduces a causal video-action model that enables robots to perform unseen manipulation tasks by following in-context human video guidance. The authors create HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks, and propose an in-context future chunk prediction objective to prevent shortcut learning. In simulation, Zero-WAM attains a 47.0% success rate on seven unseen tasks, outperforming the best video-action baseline by 29.5 percentage points, and demonstrates real‑world generalization to complex, long‑horizon, and fine‑grained tasks.

By Jiaming Zhou, Qihang Zhang, Gangwei Xu, Cunxin Fan, Yujie Zhao, Ruilin Wang, Yiming Luo, Shuai Yang, Xing Zhu, Yujun Shen, Junwei Liang, Yinghao Xu
Hugging Face Trending Papers
Jun 2

VLESA: Vision-Language Embodied Safety Agent for Human Activity Monitoring

As AI systems increasingly assist humans in physical tasks, ensuring safety becomes paramount -- physical actions carry immediate and irreversible consequences that digital errors do not. We introduce the Vision-Language Embodied Safety Agent (VLESA), a framework that monitors human activities from egocentric video and triggers real-time safety interventions when dangerous actions are predicted.

arXiv Computer Vision
Aug 26

SeMoCo: A Semantic-First Motion Codec for Motion Language Modeling

arXiv:2608.24334v1 Announce Type: new Abstract: Discrete motion representations have substantially advanced autoregressive text-to-motion generation. However, most motion tokenizers are optimized for...

By Tianlv Huang, Hetian Guo, Ziyi Cai, Song Wang, Yanping Zhang, Zipei Fan, Xuan Song, Guangming Wu, Xin Zheng
arXiv Computer Vision
Sep 4

WorldReward: Reward Modeling for Camera-Conditioned World Models

WorldReward introduces a vision‑language model–based reward system for camera‑conditioned world models, combining action consistency and visual quality evaluation. It processes paired videos by splitting them into action‑aligned chunks, structuring visual evidence, and aggregating decisions through voting. The model is trained on a large, reasoning‑augmented preference dataset and outperforms GPT‑5.5 on a human‑annotated benchmark, improving both action execution and visual quality when applied to RL post‑training.

By Yibin Wang, Zehan Wang, Junshu Tang, Zhimin Li, Yujie Zhou, Jiazi Bu, Pengyang Ling, Feng Han, Zhixiong Zhang, Long Xing, Shengyuan Ding, Ziang Li, Cheng Jin, Yuhang Zang, Jiaqi Wang, Tianyu Pang
arXiv Computer Vision
Sep 17

Occluded Gait Recognition with Mixture of Experts: An Action Detection Perspective

The paper introduces GaitMoE, an action‑detection based mixture‑of‑experts framework for occluded gait recognition, leveraging temporal and action experts to infer missing body parts from adjacent frames and gait cycles. It also presents a new Occluded Gait database (OccGait) with diverse occlusion scenarios and annotations, and demonstrates superior performance on OccGait, OccCASIA‑B, Gait3D, and GREW datasets.

By Panjian Huang, Yunjie Peng, Saihui Hou, Chunshui Cao, Xu Liu, Zhiqiang He, Yongzhen Huang
arXiv AI
3d ago

Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models

arXiv:2609.40219v1 Announce Type: cross Abstract: World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experienc...

By Qi Lyu, Jiahua Dong, Hao Shen, Xudong Wang, Hongyuan Yu, Baichen Liu, Henghui Ding, Zhi Han, Nicu Sebe, Ivan Laptev, Fahad Shahbaz Khan, Salman Khan
arXiv AI
Jun 2

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data

arXiv:2606. 00054v1 Announce Type: cross Abstract: Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models.

By Zhiyuan Feng, Qixiu Li, Huizhi Liang, Rushuai Yang, Yichao Shen, Zhiying Du, Zhaowei Zhang, Yu Deng, Li Zhao, Hao Zhao, Zongqing Lu, Oier Mees, Marc Pollefeys, Jiaolong Yang, Baining Guo
arXiv Computer Vision
Aug 27

Skeleton-based Zero-Shot Spatio-Temporal Action Localization via Weakly-Supervised Pretraining

The paper introduces Skeleton-Language feature Pooling Switching, a weakly‑supervised vision‑language pretraining strategy for skeleton‑based zero‑shot spatio‑temporal action localization. It replaces video‑level pooling with instance‑level feature computation during inference, enabling the model to estimate unseen actions without costly annotations. Additionally, Scene‑Mixed Discriminative Contrastive Learning is proposed to separate actions at the instance level within mixed scenes using a MIL framework, and experiments on four public datasets confirm the method’s effectiveness.

By Koshiro Nagano, Fumiaki Sato, Ryo Hachiuma, Kazuki Tsutsukawa, Taiki Sekii
arXiv Machine Learning
Jun 3

VLESA: Vision-Language Embodied Safety Agent for Human Activity Monitoring

arXiv:2606. 03954v1 Announce Type: cross Abstract: As AI systems increasingly assist humans in physical tasks, ensuring safety becomes paramount -- physical actions carry immediate and irreversible consequences that digital errors do not.

By Hanjiang Hu, Yiyuan Pan, Jiaxing Li, Xusheng Luo, Alexander Robey, Na Li, Yebin Wang, Changliu Liu