arXiv Computer Vision
1d ago

Lightweight and Resource-Efficient Perception for Robotic Guide Dogs

The paper examines how multi‑camera streaming perception systems perform on heterogeneous edge platforms that share resources with other workloads. Using two end‑to‑end pipelines on a single GPU–NPU platform, the authors show that isolated single‑stream evaluations can mislead deployment decisions: while the GPU pipeline appears superior in isolation, GPU‑local contention causes deadline misses that make detections stale and can reverse the preferred placement. The study finds that the NPU pipeline, though less accurate for small and medium objects, nearly matches the GPU on large objects, and that under high contention the best placement shifts from All‑GPU to All‑NPU, achieving a 5.2× improvement in worst‑stream sAP. The authors argue that evaluation metrics should include contention sweeps, deadline‑miss rates, and worst‑stream sAP in addition to mean sAP to capture severe single‑stream degradation.

By Jinse Kwon, Yoojin Lim, Choonghan Lee, Yongseung Yu, Yongin Kwon, Jemin Lee
arXiv AI
Aug 18

Efficient Block-Layer Parallel Inference for Vision-Language-Action on Hybrid Architectures

arXiv:2608. 14586v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle platforms remains difficult because they introduce both high inference latency and strong GPU-side resource pressure.

By Haibo HU, Lianming Huang, Qiao Li, Nan Guan, Chun Jason Xue