arXiv Machine Learning

Sign in the Air to Unlock: An Interface for authentication in Virtual and Augmented Reality Powered by Point-Voxel Cross-Attention Network

arXiv:2607. 01435v1 Announce Type: cross Abstract: Significant advancement of immersive technologies such as Virtual and Augmented Reality (VR/AR) and their integration into diverse aspects of modern life need authentication interfaces that are secure, intuitive, and compatible with embodied interaction.

arXiv AI
Jun 2

Interpretable Multimodal Gesture Recognition for Drone and Mobile Robot Teleoperation via Log-Likelihood Ratio Fusion

arXiv:2602. 23694v3 Announce Type: replace-cross Abstract: Human operators are still frequently exposed to hazardous environments such as disaster zones and industrial facilities, where intuitive and reliable teleoperation of mobile robots and Unmanned Aerial Vehicles (UAVs) is essential.

By Seungyeol Baek, Jaspreet Singh, Lala Shakti Swarup Ray, Hymalai Bello, Paul Lukowicz, Sungho Suh
arXiv AI
Jul 20

EgoExoMoCap: Distributed Ego-Exo Human Motion Capture

arXiv:2607. 15868v1 Announce Type: cross Abstract: Human motion capture from head-mounted devices (HMDs) offers a scalable way to acquire real-world human motion and interaction data, which is crucial for applications in embodied AI and VR/AR.

By Jiaxi Jiang, Bharat Lal Bhatnagar, Nan Yang, Lingni Ma, Sebastian Starke, Robin Kips, Nadine Bertsch, Christian Holz, Federica Bogo
arXiv AI
Jun 6

Tamaththul3D: High-Fidelity 3D Saudi Sign Language Avatars from Monocular Video

arXiv:2605. 05367v2 Announce Type: replace-cross Abstract: Existing 3D sign language avatar reconstruction methods are developed and evaluated exclusively on Western sign languages, and no 3D parametric annotations exist for any Arabic Sign Language dataset, a gap that blocks the development of avatar-based accessibility applications for the Arab Deaf community.

By Eyad Alghamdi, Sattam Altuuaim, Obay Ghulam, Abdulrahman Qutah, Yousef Basoodan
arXiv AI
Sep 1

Toward Postural State Classification in Immersive VR with Multimodal Data and Explainability Analysis

This study evaluates machine learning and deep‑learning models for classifying balanced versus imbalanced postural states in immersive virtual reality using a multimodal dataset of kinematic, EMG, and EDA signals. The Mamba‑inspired CNN (MI‑CNN) achieved the highest accuracy (96.76%) and, through SHapley Additive exPlanations (SHAP), identified kinematic features as the most influential for detecting imbalance. Even after reducing input dimensionality by 33% based on SHAP importance, the model maintained near‑optimal performance (0.957 accuracy and F1‑score).

By Nipa Anjum, Md Irfan Pavel, Robert Gonzalez Jr, Kevin Desai, Alberto Cordova, M. Rasel Mahmud, John Quarles
arXiv Computer Vision
Sep 25

PHOSA: Photorealistic 3D Sign Avatar Modeling and Benchmark

PHOSA introduces MVSign, the first multi‑view Chinese sign language dataset co‑designed with Deaf experts, featuring diverse gestures and rich annotations. The authors develop a hybrid fitting pipeline for accurate SMPL‑X annotation and propose a decoupled sign avatar representation that isolates body, head, and hand components, coupled with a motion‑aware sampling strategy to handle motion blur and balance gesture diversity. Experiments show high‑fidelity visual results on MVSign, especially in detailed hand and facial regions, and good generalization to in‑the‑wild monocular sign language videos.

By Haodong Wang, Hezhen Hu, Wengang Zhou, Houqiang Li
arXiv Machine Learning
Sep 15

Speak to the City: Multimodal Resolution for Outside-the-Vehicle References

arXiv:2609.14691v1 Announce Type: cross Abstract: As autonomous vehicles and Extended Reality (XR) headsets enable novel in-car interactions, seamlessly querying physical landmarks, known as Outside-...

By Alireza Parchami (Mercedes-Benz Tech Innovation GmbH, Saarland University), Artin Saberpour (Saarland University), Robin Connor Schramm (Mercedes-Benz Tech Innovation GmbH, RheinMain University of Applied Sciences), J\"urgen Steimle (Saarland University), Ulrich Schwanecke (RheinMain University of Applied Sciences)
arXiv Machine Learning
Aug 18

Helios 2.0: A Robust, Ultra-Low Power Gesture Recognition System Optimised for Event-Sensor based Wearables

arXiv:2503. 07825v3 Announce Type: replace-cross Abstract: We present an advance in wearable technology: a mobile-optimized, real-time, ultra-low-power event camera system that enables natural hand gesture control for smart glasses, dramatically improving user experience.

By Prarthana Bhattacharyya, Joshua Mitton, Ryan Page, Owen Morgan, Oliver Powell, Benjamin Menzies, Gabriel Homewood, Kemi Jacobs, Paolo Baesso, Taru Muhonen, Richard Vigars, Louis Berridge
arXiv Computer Vision
Aug 21

DreamHand: Repurposing Video Diffusion Models for Occlusion-Robust Egocentric 3D Hand Motion Recovery

arXiv:2608. 20308v1 Announce Type: new Abstract: Egocentric video offers scalable manipulation data for embodied AI, yet recovering metric 3D hand trajectories remains challenging due to severe object occlusion and frequent out-of-sight gaps.

By Yufei Liu, Xixi Wang, Hao Li, Ganlong Zhao, Kaitong Cai, Chengkai Jin, Chunxiao Liu, Jianbo Liu, Siyuan Huang, Xingang Pan, Hongsheng Li