arXiv Computer Vision By Jinse Kwon, Yoojin Lim, Choonghan Lee, Yongseung Yu, Yongin Kwon, Jemin Lee

Lightweight and Resource-Efficient Perception for Robotic Guide Dogs

Read the original on arXiv Computer Vision →

The paper examines how multi‑camera streaming perception systems perform on heterogeneous edge platforms that share resources with other workloads. Using two end‑to‑end pipelines on a single GPU–NPU platform, the authors show that isolated single‑stream evaluations can mislead deployment decisions: while the GPU pipeline appears superior in isolation, GPU‑local contention causes deadline misses that make detections stale and can reverse the preferred placement. The study finds that the NPU pipeline, though less accurate for small and medium objects, nearly matches the GPU on large objects, and that under high contention the best placement shifts from All‑GPU to All‑NPU, achieving a 5.2× improvement in worst‑stream sAP. The authors argue that evaluation metrics should include contention sweeps, deadline‑miss rates, and worst‑stream sAP in addition to mean sAP to capture severe single‑stream degradation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Machine Learning
Sep 14

Efficient Vision-Language-Action Management and Serving for Robot Factories

Robion is a new serving and management system designed to run Vision‑Language‑Action (VLA) models on multi‑GPU edge servers for robot factories. It splits the VLM and ADiT stages within a single GPU, shares streams across multiple models, and prioritizes requests by remaining SLO time, enabling high robot load while meeting strict latency requirements. In experiments, Robion achieves 6.7× higher robot load than vLLM‑Omni and 1.5× higher than a monolithic pipeline, and can serve 64 robots on a 4‑GPU server with 98% SLO attainment.

By Dionysios Adamopoulos, Nattapol Chanpaisit, Basel Fakhri, Christina Giannoula
arXiv AI
Aug 18

Efficient Block-Layer Parallel Inference for Vision-Language-Action on Hybrid Architectures

arXiv:2608. 14586v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle platforms remains difficult because they introduce both high inference latency and strong GPU-side resource pressure.

By Haibo HU, Lianming Huang, Qiao Li, Nan Guan, Chun Jason Xue
arXiv AI
Sep 17

Visual Perception Engine: Fast and Flexible Multi-Head Inference for Robotic Vision Tasks

Visual Perception Engine (VPEngine) is a modular framework that enables efficient GPU usage for robotic vision tasks by sharing a foundation model backbone across multiple specialized task heads. It eliminates redundant feature extraction, supports dynamic task prioritization, and achieves up to 3× speedup over sequential execution. The open‑source Python implementation, with ROS2 C++ bindings, delivers real‑time performance (≥50 Hz) on NVIDIA Jetson Orin AGX using TensorRT‑optimized models.

By Jakub {\L}ucki, Jonathan Becktor, Georgios Georgakis, Rob Royce, Shehryar Khattak