arXiv Machine Learning

Lightweight Neural Networks for Affordance Segmentation: Enhancement of the Decoder Module

arXiv:2607. 29473v1 Announce Type: cross Abstract: The deployment of deep neural networks for visual affordance segmentation on wearable robots poses may prove critical, due to some conflicting aspects of the problem.

arXiv AI
4d ago

UniAfford: Token-Routed Multitask Learning for Generalizable 2D-3D Affordance Perception

arXiv:2609.37264v1 Announce Type: cross Abstract: Affordance perception aims to localize actionable regions supporting embodied interaction, yet 2D and 3D affordance grounding have evolved as separat...

By Yuhao Liu, Yiming Zhong, Hanqing Wang, Shaocheng Yan, Yuhang Zhang, Wenzhou Lyu, Ziyang Ding, Wei Zhang, Xue Zhao, Jin Pan, Yuexin Ma, Xinge Zhu
arXiv Computer Vision
Sep 21

Recursive Block-Diagonal Coupling for Resource-Efficient Training of Vision Models

The paper introduces Recursive Block-Diagonal Coupling (RBDC), a training protocol that builds wide vision models by recursively coupling narrower, independently trained models in a parameter‑free block‑diagonal manner. RBDC allows flexible allocation of training budgets across all models and, when applied to vision transformers (DeiT) and convolutional networks (ResNet) on ImageNet, achieves a 30% reduction in FLOPs while maintaining similar test accuracies. Additionally, models trained with RBDC outperform those from existing growth methods at the same training FLOPs and serve as stronger backbones for downstream tasks such as object detection and instance segmentation.

By Maxim Henry, Adrien Deli\`ege, S\'ebastien Pi\'erard, Marc Van Droogenbroeck
arXiv AI
6d ago

Models Got Talent: Identifying High Performing Wearable Human Activity Recognition Models Without Training

The paper investigates the use of Zero Cost Proxies (ZCPs) to identify high‑performing wearable Human Activity Recognition (HAR) models without full training. Eight ZCPs were evaluated across six benchmark HAR datasets, showing that the top‑predicted architectures achieve performance within 7% of fully trained models, and training the top‑10 predictions reaches within 2% of full training. This demonstrates that ZCPs can significantly reduce computational costs while maintaining competitive accuracy in sensor‑based HAR tasks.

By Richard Goldman, Varun Komperla, Thomas Ploetz, Harish Haresamudram
arXiv AI
Sep 17

Visual Perception Engine: Fast and Flexible Multi-Head Inference for Robotic Vision Tasks

Visual Perception Engine (VPEngine) is a modular framework that enables efficient GPU usage for robotic vision tasks by sharing a foundation model backbone across multiple specialized task heads. It eliminates redundant feature extraction, supports dynamic task prioritization, and achieves up to 3× speedup over sequential execution. The open‑source Python implementation, with ROS2 C++ bindings, delivers real‑time performance (≥50 Hz) on NVIDIA Jetson Orin AGX using TensorRT‑optimized models.

By Jakub {\L}ucki, Jonathan Becktor, Georgios Georgakis, Rob Royce, Shehryar Khattak
arXiv AI
Jun 30

Tactile Gesture Recognition with Built-in Joint Sensors for Industrial Robots

arXiv:2508. 12435v2 Announce Type: replace-cross Abstract: While gesture recognition using vision or robot skins is an active research area in Human-Robot Collaboration (HRC), this paper explores deep learning methods relying solely on a robot's built-in joint sensors, eliminating the need for external sensors.

By Deqing Song, Weimin Yang, Maryam Rezayati, Hans Wernher van de Venn