arXiv:2607. 28293v1 Announce Type: cross Abstract: While depth sensors have the potential to complement RGB data for affordance segmentation in wearable robots, their usage seems to remain underexplored.
By Edoardo Ragusa, Giovanni Paolo Canuti, Simone Lugani, Rodolfo Zunino, Paolo Gastaldo
arXiv:2605.03927v3 Announce Type: replace
Abstract: Vision-language models have demonstrated strong performance across robotic perception and instruction-following tasks. However, they still struggle...
By Xiaowen Sun, Matthias Kerzel, Mengdi Li, Xufeng Zhao, Paul Striker, Stefan Wermter
arXiv:2607. 16222v1 Announce Type: cross Abstract: This paper presents ARGO, a smart eyewear platform designed to bridge ergonomic comfort, high computational throughput, and energy efficiency.
By Andrea Giudici, Christian Veronesi, Pietro Bartoli, Mario Cali\`o, Aurelio Teliti, Giacomo Gervasoni, Diana Trojaniello, Franco Zappa
arXiv:2609.37264v1 Announce Type: cross
Abstract: Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separat...
By Yuhao Liu, Yiming Zhong, Hanqing Wang, Shaocheng Yan, Yuhang Zhang, Wenzhou Lyu, Ziyang Ding, Wei Zhang, Xue Zhao, Jin Pan, Yuexin Ma, Xinge Zhu
arXiv:2609.10018v1 Announce Type: new
Abstract: EdgeAI systems are increasingly employing computer vision applications to enable intelligent, on-device decision-making in real-time. However, these de...
By Sudaksh Kalra, Dolly Sapra
arXiv:2606. 30632v1 Announce Type: cross Abstract: Can the robot use a plate to cut a cake if no knife is available?
By Yuhong Deng, Yuyao Liu, David Hsu
arXiv:2607. 05568v1 Announce Type: cross Abstract: Representing 3D shapes as compact sets of geometric primitives is fundamental to robotics, simulation, and scene understanding.
By Gregor Kobsik, Tim Elsner, Leif Kobbelt
The paper introduces Recursive Block-Diagonal Coupling (RBDC), a training protocol that builds wide vision models by recursively coupling narrower, independently trained models in a parameter‑free block‑diagonal manner. RBDC allows flexible allocation of training budgets across all models and, when applied to vision transformers (DeiT) and convolutional networks (ResNet) on ImageNet, achieves a 30% reduction in FLOPs while maintaining similar test accuracies. Additionally, models trained with RBDC outperform those from existing growth methods at the same training FLOPs and serve as stronger backbones for downstream tasks such as object detection and instance segmentation.
By Maxim Henry, Adrien Deli\`ege, S\'ebastien Pi\'erard, Marc Van Droogenbroeck
The paper investigates the use of Zero Cost Proxies (ZCPs) to identify high‑performing wearable Human Activity Recognition (HAR) models without full training. Eight ZCPs were evaluated across six benchmark HAR datasets, showing that the top‑predicted architectures achieve performance within 7% of fully trained models, and training the top‑10 predictions reaches within 2% of full training. This demonstrates that ZCPs can significantly reduce computational costs while maintaining competitive accuracy in sensor‑based HAR tasks.
By Richard Goldman, Varun Komperla, Thomas Ploetz, Harish Haresamudram
Visual Perception Engine (VPEngine) is a modular framework that enables efficient GPU usage for robotic vision tasks by sharing a foundation model backbone across multiple specialized task heads. It eliminates redundant feature extraction, supports dynamic task prioritization, and achieves up to 3× speedup over sequential execution. The open‑source Python implementation, with ROS2 C++ bindings, delivers real‑time performance (≥50 Hz) on NVIDIA Jetson Orin AGX using TensorRT‑optimized models.
By Jakub {\L}ucki, Jonathan Becktor, Georgios Georgakis, Rob Royce, Shehryar Khattak
arXiv:2508. 12435v2 Announce Type: replace-cross Abstract: While gesture recognition using vision or robot skins is an active research area in Human-Robot Collaboration (HRC), this paper explores deep learning methods relying solely on a robot's built-in joint sensors, eliminating the need for external sensors.
By Deqing Song, Weimin Yang, Maryam Rezayati, Hans Wernher van de Venn
arXiv:2604.00634v3 Announce Type: replace-cross
Abstract: Panoptic segmentation is a key enabler for robotic perception, as it unifies semantic understanding with object-level reasoning. However, the...
By Calvin Galagain, Martyna Poreba, Fran\c{c}ois Goulette, Cyrill Stachniss