arXiv:2606. 31127v1 Announce Type: cross Abstract: To enable personalized, real-time coaching using Augmented Reality glasses or fixed camera setups in domains such as sports, cooking, or music, a system must understand not just what a person does, but how well they execute an activity.
By Bj\"orn Braun, Christian Holz
arXiv:2608. 05115v1 Announce Type: cross Abstract: Can computer vision help make classrooms safer?
By Paritosh Parmar, Landy Lan, Hong Yang, Chen Yi, Chiat Pin Tay
arXiv:2609.36502v1 Announce Type: cross
Abstract: Past research using log data has faced the "learning system wall," whereby few methods exist for generalizing models of student learning across platf...
By Danielle R. Thomas, Marie Cynthia Abijuru Kamikazi, Ashish Gurung, Ishan Miglani, Shivang Gupta, Zachary Levonian, Conrad Borchers, Kenneth R. Koedinger
Can computer vision help make classrooms safer? In this pilot study, we investigate privacy-aware and computationally efficient classroom incident recognition from CCTV-style observations.
arXiv:2608. 16222v1 Announce Type: cross Abstract: Humanoid intelligence requires learning over an extremely diverse space of whole-body motions and physically grounded interactions.
By Jiahao Ji, Ji Ma, Runhan Zhang, Runyi Yu, Wenjia Wang, Weiheng Chi, Qianqian Peng, Weichao Yan, Yongfei Gu, Ye Tian, Ting Wu, Longwei Li, Chun Yuan, Ruoli Dai, Lei Han
MyoMechanix is a multimodal dataset and framework for action quality assessment that incorporates muscle activity and other physiological signals alongside visual data. It contains over 7,500 samples of 20 weight‑loaded actions from 38 subjects, with synchronized RGB video, 3D pose, sEMG, and additional signals. The accompanying Fitness Knowledge Graph structures expert annotations into relationships among actions, phases, key steps, errors, and corrective feedback, enabling compositional scoring and interpretable assessment through the CUBIST engine. The project also introduces MyoMechanix‑AQA, MyoMechanix‑VideoQA, and a novel MyoMechanix‑Video2EMG task, demonstrating that multimodal sensing and structured representations improve performance, interpretability, and error attribution.
By Hao Yin, Paritosh Parmar, Lijun Gu, Lin Xu, Tianxiao Guo, Xiujin Liu, Tianyou Zheng, Yang Zhang, Weiwei Fu
arXiv:2609.09300v1 Announce Type: new
Abstract: Video understanding demands a convergence of complementary capabilities across perception, temporal understanding, and complex reasoning, which are dif...
By Zhenxin Qin, Peng Shi, Cong Han, Yinlong Qian, Zequn Jie, Lin Ma
arXiv:2608. 15861v1 Announce Type: new Abstract: Fine-grained wrist activity recognition can support applications such as procedural step guidance and context-aware assistance, yet acquiring labeled data for every new task, user, and activity granularity remains a bottleneck.
By Aidan Bradshaw, Riku Arakawa, Xin Liu, Karan Ahuja
arXiv:2202. 14019v3 Announce Type: replace-cross Abstract: Maintaining proper form while exercising is important for preventing injuries and maximizing muscle mass gains.
By Paritosh Parmar, Amol Gharat, Helge Rhodin
arXiv:2407. 13053v2 Announce Type: replace-cross Abstract: Digital textbook (e-book) systems record student interactions with textbooks as a sequence of events called EventStream data.
By Yuma Miyazaki, Valdemar \v{S}v\'abensk\'y, Yuta Taniguchi, Fumiya Okubo, Tsubasa Minematsu, Atsushi Shimada
arXiv:2608.27562v1 Announce Type: new
Abstract: Translating continuous, noisy egocentric video streams into discrete, temporally ordered action steps is fraught with visual challenges. Heavy ego-moti...
By Anubhav Gupta, Archit Kambhamettu, Vatsal Agarwal, Pulkit Kumar, Abhinav Shrivastava
arXiv:2608.08273v2 Announce Type: replace-cross
Abstract: Vision-based embodied agents executing multi-step natural language instructions require feedback mechanisms that assess task progress over co...
By Hwanhee Kim, Jaehyun Jang, Seungmin Cha, Hyeonseo Yun, Donghoon Lee, Chang D. Yoo