arXiv AI

Interpretable Multimodal Gesture Recognition for Drone and Mobile Robot Teleoperation via Log-Likelihood Ratio Fusion

arXiv:2602. 23694v3 Announce Type: replace-cross Abstract: Human operators are still frequently exposed to hazardous environments such as disaster zones and industrial facilities, where intuitive and reliable teleoperation of mobile robots and Unmanned Aerial Vehicles (UAVs) is essential.

arXiv AI
Jun 2

DeepIPCv3: Event-Aware Multi-Modal Sensor Fusion for Sudden Pedestrian Crossing Avoidance

arXiv:2606. 01277v1 Announce Type: cross Abstract: Current end-to-end autonomous driving systems predominantly rely on frame-based sensors, which suffer from inherent perception latency and motion blur during highly dynamic encounters, specifically sudden pedestrian crossings.

By Oskar Natan, Andi Dharmawan, Aufaclav Zatu Kusuma Frisky, Jazi Eko Istiyanto, Jun Miura
arXiv AI
Jul 20

EgoExoMoCap: Distributed Ego-Exo Human Motion Capture

arXiv:2607. 15868v1 Announce Type: cross Abstract: Human motion capture from head-mounted devices (HMDs) offers a scalable way to acquire real-world human motion and interaction data, which is crucial for applications in embodied AI and VR/AR.

By Jiaxi Jiang, Bharat Lal Bhatnagar, Nan Yang, Lingni Ma, Sebastian Starke, Robin Kips, Nadine Bertsch, Christian Holz, Federica Bogo
arXiv Machine Learning
Aug 18

MiDAS: A Multimodal Data Acquisition System and Dataset for Robot-Assisted Minimally Invasive Surgery

arXiv:2602. 12407v3 Announce Type: replace-cross Abstract: Background: Robot-assisted minimally invasive surgery (RMIS) research increasingly relies on multimodal data, yet access to proprietary robot telemetry remains a major barrier.

By Keshara Weerasinghe (MD), Seyed Hamid Reza Roodabeh (MD), Andrew Hawkins (MD), Zhaomeng Zhang, Zachary Schrader, Homa Alemzadeh
arXiv Computer Vision
Sep 16

EventEgoHands++: Event-based Egocentric 3D Hand Mesh Reconstruction with Real Dataset

EventEgoHands++ is a new framework for reconstructing 3D hand meshes from egocentric event-based cameras. It introduces a Hand Detector that provides instance-level bounding boxes and masks for left and right hands, and an Adaptive Attention module that uses these detections to model spatial relationships and interactions. The authors extend the synthetic N-HOT3D dataset and create EEH‑R, a large real-world event-based egocentric hand dataset with about 1 million annotated frames, and show that their method outperforms existing baselines on both synthetic and real data.

By Ryosei Hara, Wataru Ikeda, Masashi Hatano, Mariko Isogawa
arXiv Computer Vision
Sep 7

SocioGesture: Real-Time and Adaptive Social Gesture Perception for Human-Robot Interaction

SocioGesture is a real‑time, adaptive system for recognizing social gestures in human‑robot interaction. It employs a compact, confidence‑aware body‑hand skeleton representation and a lightweight dual‑stream model that fuses body motion with hand articulation, enabling low‑latency onboard recognition. The model is trained with occlusion‑aware skeleton corruption to handle missing hands, occluded arms, and unstable keypoints, and it can expand its gesture vocabulary during deployment by saving uncertain interaction segments for offline labeling.

By Wenjin Fu, Li-Fan Wu, Jerin Peter, Chip Huyen, Boyuan Chen, Jan Liphardt
arXiv Machine Learning
Jun 29

Cross-view Multimodal Vision-Based Assessment Framework for Traditional Chinese Medicine Rehabilitation Training

arXiv:2606. 28104v1 Announce Type: cross Abstract: Vision-based assessment can provide convenient and cost-effective evaluation in Traditional Chinese Medicine (TCM) rehabilitation training, where action quality assessment (AQA) from computer vision offers a promising solution.

By Francis Xiatian Zhang, Hao Yao, Shengxuan Chen, Hong Zhu, Hongxiao Jia, Sisi Zheng, Hubert P. H. Shum