MEgoVista is an offline pipeline that converts a single unprepared egocentric video into metric two‑hand and head motion within a gravity‑aligned world frame. It uniquely reconstructs motion in environments beyond studio volumes, uses calibrated stereo for absolute scale, and evaluates its outputs against independent optical capture to audit accuracy. The system thus expands the settings where high‑fidelity hand‑motion labels can be generated from natural, head‑worn recordings.
By Jiangong Xiao (Northwestern Polytechnical University), Zhihao Zhang (Xi'an Jiaotong University), Yifei Dong (Maniformer), Chao Ma (Maniformer), Zhouyi Jin (Maniformer), Zhiwen Hou (Maniformer), Li Liu (Maniformer), Weihuang Chen (Xi'an Jiaotong University), Hongbin Sun (Xi'an Jiaotong University), Maoqing Yao (Maniformer)
arXiv:2609.18772v1 Announce Type: new
Abstract: Sign language processing advances rapidly for high-resource languages such as American Sign Language (ASL), yet most of the world's sign languages lack...
By Marcel Granero-Moya, Carolina del Corral Farrar\'os, Gloria Haro, Coloma Ballester, Ricardo Marques
We present the first experiments on isolated sign language recognition (ISLR) for Icelandic Sign Language (ÍTM). We use ÍTM SignWiki, a dataset derived from a bilingual Icelandic--ÍTM online dictionar...
arXiv:2609.24424v1 Announce Type: new
Abstract: Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict...
By Kaiwen Ren, Yiran Jiang, Yongjing Ye, Shihong Xia
arXiv:2609.25862v1 Announce Type: new
Abstract: We present the first experiments on isolated sign language recognition (ISLR) for Icelandic Sign Language (\'ITM). We use \'ITM SignWiki, a dataset der...
By Finnur \'Ag\'ust Ingimundarson, Gu{\dh}n\'y Bj\"ork {\TH}orvaldsd\'ottir, Mathias M\"uller, Sarah Ebling
arXiv:2608. 03127v1 Announce Type: cross Abstract: Hand motion carries the finest-grained information in human activity, yet the representations behind hand generation, understanding, and robot learning are overwhelmingly continuous--joint angles or MANO parameters.
By Haoyu Gu, Haotian Lu, Jingrun Du, Xiao-Ping Zhang
The paper argues that in structured image classification, the key question is what information a classifier should receive before making a prediction, rather than which algorithm performs best. It proposes a systematic framework for constructing landmark-derived representations—such as coordinate, distance, angle, and hybrid features—and evaluates them on static hand gesture recognition. Experiments show that hybrid representations, which combine complementary geometric components, outperform raw coordinate features and other single-type representations, highlighting the importance of thoughtful feature construction.
By Saravana Mauree, Sakshi Arya
The paper presents a training‑free method for detecting which holds a climber uses in sport climbing videos by leveraging a frozen foundation pose model (Sapiens) that provides fingertip and toe keypoints. Using a simple proximity test, mutual exclusion, and a temporal‑persistence rule, the approach achieves high F_1 scores (up to 90.2%) on the Way Up dataset without any climbing‑specific training, outperforming repurposed pose pipelines. The resulting automatic predictions enable accurate coaching statistics, such as climb time and pace, with Pearson correlations of 1.00 and 0.94 respectively.
By Abu Bakar, Abdullah Aftab, Amir Hamza
Hand Pose Estimation (HPE) is a fundamental technology for various applications such as AR/VR and robotics. In these applications, the visibility of each hand joint in the image is crucial for assessing the reliability of estimation results under occlusion.
arXiv:2608. 10588v1 Announce Type: cross Abstract: Purpose: Fine-grained handshape recognition supports computational sign-language transcription, recognition, and translation, but broad, phonetically defined visual inventories with signer-aware evaluation remain limited.
By Ushnish Sarkar, Suvajit Patra, Bhaswar Chattopadhyay, Pranab Singha Roy, Tapas Samanta
arXiv:2609.14122v1 Announce Type: new
Abstract: We study the challenge of sign language video mimicking: given a driving video and a single reference frame, synthesize a video where the target signer...
By Zhewen He (New York University Abu Dhabi), Junyi Yu (New York University Abu Dhabi), Haomian Huang (New York University Abu Dhabi), Zhenhua Li (ChatSign Technology), Yi Fang (New York University Abu Dhabi, ChatSign Technology)
EventEgoHands++ is a new framework for reconstructing 3D hand meshes from egocentric event-based cameras. It introduces a Hand Detector that provides instance-level bounding boxes and masks for left and right hands, and an Adaptive Attention module that uses these detections to model spatial relationships and interactions. The authors extend the synthetic N-HOT3D dataset and create EEH‑R, a large real-world event-based egocentric hand dataset with about 1 million annotated frames, and show that their method outperforms existing baselines on both synthetic and real data.
By Ryosei Hara, Wataru Ikeda, Masashi Hatano, Mariko Isogawa