arXiv:2606. 17615v1 Announce Type: cross Abstract: Estimating human proficiency from video is a key challenge for automated skill assessment, with applications in sports coaching, music pedagogy, surgical training, and workplace learning.
By Edoardo Bianchi, Antonio Liotta
The paper introduces a redundancy-aware fusion framework for EgoExo proficiency estimation, which integrates fine-grained motion cues from egocentric views with spatial context from exocentric views. It identifies multiview redundancy and overfitting as key challenges and proposes two modules—AdaMVS for adaptive view selection and VIB-GB for compressing redundant signals—to address them. Experiments on EgoExo-4D and EgoExo-Fitness show that the method learns to select informative views and fuse them effectively, achieving state‑of‑the‑art results.
By Xu Dong, Wanqing Li, Anthony Adeyemi-Ejeye, Andrew Gilbert
arXiv:2609.23492v1 Announce Type: new
Abstract: Perception for embodied agents is video-based, often multi-view (ego, exo, or both), and inherently continual, with simultaneous task and viewpoint shi...
By Hongwei Yan, Kanglei Zhou, Yuchen Liu, Qingyu Shi, Yi Zhong, Liyuan Wang
arXiv:2606. 00054v1 Announce Type: cross Abstract: Recent progress in generalizable embodied control has been driven by large-scale pretraining of Vision-Language-Action (VLA) models.
By Zhiyuan Feng, Qixiu Li, Huizhi Liang, Rushuai Yang, Yichao Shen, Zhiying Du, Zhaowei Zhang, Yu Deng, Li Zhao, Hao Zhao, Zongqing Lu, Oier Mees, Marc Pollefeys, Jiaolong Yang, Baining Guo
The paper introduces GaitMoE, an action‑detection based mixture‑of‑experts framework for occluded gait recognition, leveraging temporal and action experts to infer missing body parts from adjacent frames and gait cycles. It also presents a new Occluded Gait database (OccGait) with diverse occlusion scenarios and annotations, and demonstrates superior performance on OccGait, OccCASIA‑B, Gait3D, and GREW datasets.
By Panjian Huang, Yunjie Peng, Saihui Hou, Chunshui Cao, Xu Liu, Zhiqiang He, Yongzhen Huang
CLAP is a cross-embodiment framework for action‑conditioned video generation that can be trained on diverse internet‑scale videos from both humans and robots. It reconciles different action spaces—end‑effector poses, language instructions, and latent actions—using a curriculum that first learns physics priors from unlabeled video and then grounds them in real‑world action spaces for zero‑shot deployment. The resulting models match or exceed state‑of‑the‑art single‑embodiment models in challenging environments and support few‑shot adaptation across a wide range of robot morphologies.
By Kechen Liu, Ola Shorinwa