arXiv:2602. 23694v3 Announce Type: replace-cross Abstract: Human operators are still frequently exposed to hazardous environments such as disaster zones and industrial facilities, where intuitive and reliable teleoperation of mobile robots and Unmanned Aerial Vehicles (UAVs) is essential.
By Seungyeol Baek, Jaspreet Singh, Lala Shakti Swarup Ray, Hymalai Bello, Paul Lukowicz, Sungho Suh
Perceiving human motion and intent at long range is a prerequisite for socially intelligent aerial robots, yet the data to learn it barely exists. We introduce Drones2BodyLanguage, a dataset grounding human motion in real UAV footage: avatars manifesting ten communicative intents are placed into unmodified 4K drone scenes with metrically correct position, scale and orientation, maintained over hundreds of frames of camera motion.
arXiv:2608.16081v2 Announce Type: replace
Abstract: Open-weight and frontier vision-language models (VLMs) perform well on general image understanding, but their ability to interpret fine-grained han...
By Taegang Kim, Saleh Afroogh, Junfeng Jiao
arXiv:2607. 01435v1 Announce Type: cross Abstract: Significant advancement of immersive technologies such as Virtual and Augmented Reality (VR/AR) and their integration into diverse aspects of modern life need authentication interfaces that are secure, intuitive, and compatible with embodied interaction.
By Neda Abdolrahimi, Thiru Siddharth, Frank Sicongchen, Vir V Phoha
arXiv:2607. 14675v1 Announce Type: cross Abstract: Robust human-robot interaction in complex environments requires accurate gesture perception, semantic scene understanding, and reliable task planning under limited onboard computing resources.
By Zihan Guo, Xiaoqi Li
PHOSA introduces MVSign, the first multi‑view Chinese sign language dataset co‑designed with Deaf experts, featuring diverse gestures and rich annotations. The authors develop a hybrid fitting pipeline for accurate SMPL‑X annotation and propose a decoupled sign avatar representation that isolates body, head, and hand components, coupled with a motion‑aware sampling strategy to handle motion blur and balance gesture diversity. Experiments show high‑fidelity visual results on MVSign, especially in detailed hand and facial regions, and good generalization to in‑the‑wild monocular sign language videos.
By Haodong Wang, Hezhen Hu, Wengang Zhou, Houqiang Li
arXiv:2605. 05367v2 Announce Type: replace-cross Abstract: Existing 3D sign language avatar reconstruction methods are developed and evaluated exclusively on Western sign languages, and no 3D parametric annotations exist for any Arabic Sign Language dataset, a gap that blocks the development of avatar-based accessibility applications for the Arab Deaf community.
By Eyad Alghamdi, Sattam Altuuaim, Obay Ghulam, Abdulrahman Qutah, Yousef Basoodan
arXiv:2606. 04708v1 Announce Type: cross Abstract: Universal Manipulation Interface (UMI) enables scalable real-world robot data collection without hardware-specific teleoperation, yet leveraging UMI data to train large-scale Vision-Language-Action (VLA) models remains fundamentally challenging.
By Siyuan Yang, Linzheng Guo, Ouyang Lu, Zhaxizhuoma, Daoran Zhang, Xinmiao Wang, Ting Xiao, Fangzheng Yan, Zhijun Chen, Yan Ding, Chao Yu, Chenjia Bai, Xuelong Li
arXiv:2602. 20958v2 Announce Type: replace-cross Abstract: Vision-based Unmanned Aerial Vehicles (UAVs) frameworks aid human search tasks by detecting and recognizing specific individuals, then tracking and following them while maintaining a safe distance.
By Luka \v{S}iktar, Branimir \'Caran, Bojan \v{S}ekoranja, Marko \v{S}vaco
arXiv:2608.27997v1 Announce Type: new
Abstract: Air-ground cross-view referring person detection is a necessary component in the language-to-perception-to-control chain of collective embodied intelli...
By Zhoupeng Guo, Xinjie Yao, Yunqi Zhu, Zhihe Fan, Siqi Zhao, Jianjun Chen, Yichen Dong, Yan Fan, Pengfei Zhu
SocioGesture is a real‑time, adaptive system for recognizing social gestures in human‑robot interaction. It employs a compact, confidence‑aware body‑hand skeleton representation and a lightweight dual‑stream model that fuses body motion with hand articulation, enabling low‑latency onboard recognition. The model is trained with occlusion‑aware skeleton corruption to handle missing hands, occluded arms, and unstable keypoints, and it can expand its gesture vocabulary during deployment by saving uncertain interaction segments for offline labeling.
By Wenjin Fu, Li-Fan Wu, Jerin Peter, Chip Huyen, Boyuan Chen, Jan Liphardt
arXiv:2503. 07825v3 Announce Type: replace-cross Abstract: We present an advance in wearable technology: a mobile-optimized, real-time, ultra-low-power event camera system that enables natural hand gesture control for smart glasses, dramatically improving user experience.
By Prarthana Bhattacharyya, Joshua Mitton, Ryan Page, Owen Morgan, Oliver Powell, Benjamin Menzies, Gabriel Homewood, Kemi Jacobs, Paolo Baesso, Taru Muhonen, Richard Vigars, Louis Berridge