Beyond MACs: Hardware Efficient Architecture Design for Vision Backbones
arXiv:2603. 26551v2 Announce Type: replace-cross Abstract: Vision backbone networks play a central role in modern computer vision.
Fast-BEV++ tackles the trade‑off between accuracy and efficiency in vision‑only Bird’s‑Eye‑View perception by adopting two guiding principles: Fast by Algorithm and Deployable by Design. The method decomposes view transformation into a hardware‑oriented Index‑Gather‑Reshape pipeline, removing the need for custom kernels and delivering at least a three‑fold speedup over baseline approaches. On the nuScenes benchmark, Fast‑BEV++ achieves 0.488 NDS while running in real time at over 134 FPS, with depth supervision providing consistent accuracy gains and the architecture enabling seamless deployment on production‑level platforms.
arXiv:2603. 26551v2 Announce Type: replace-cross Abstract: Vision backbone networks play a central role in modern computer vision.
arXiv:2606. 31938v1 Announce Type: cross Abstract: Deploying Vision Transformer (ViT) models on edge platforms remains challenging due to their high computational demands and the architectural heterogeneity of modern hybrid ViT models, which incorporate both fully connected and convolutional layers.
arXiv:2609.23733v1 Announce Type: new Abstract: Feed-forward visual geometry models such as the Visual Geometry Grounded Transformer (VGGT) have recently enabled direct 3D reconstruction from multi-v...
arXiv:2607. 22714v1 Announce Type: cross Abstract: Real-time perception is a foundational requirement for advanced driver assistance systems (ADAS) and autonomous vehicles, yet embedded automotive platforms impose severe constraints on compute, memory, and power.
The paper introduces a hardware‑aware framework that uses genetic programming to evolve layer‑specific scalar functions for Vision Transformers, replacing traditional LayerNorm with efficient, heterogeneous approximations. By applying a post‑training re‑alignment strategy, the method eliminates the need for full model retraining while achieving 90‑93% variance capture and recovering over 84% of ImageNet‑1K Top‑1 accuracy for ViT‑B and ViT‑L. The resulting architecture removes the global reduction bottleneck, reducing arithmetic complexity and off‑chip memory traffic, thereby enabling efficient deployment of ViTs on edge accelerators.
arXiv:2608. 11770v1 Announce Type: cross Abstract: Edge-deployed vision systems in target recognition, surveillance, autonomous vehicles, and drone domains require hierarchical inference pipelines where a detection model identifies objects of interest and downstream classifiers provide fine-grained attribute analysis.
arXiv:2609.23974v1 Announce Type: new Abstract: Foundation models are endowing autonomous systems with greater intelligence, enabling a more comprehensive understanding of the environment through vis...
Visual Perception Engine (VPEngine) is a modular framework that enables efficient GPU usage for robotic vision tasks by sharing a foundation model backbone across multiple specialized task heads. It eliminates redundant feature extraction, supports dynamic task prioritization, and achieves up to 3× speedup over sequential execution. The open‑source Python implementation, with ROS2 C++ bindings, delivers real‑time performance (≥50 Hz) on NVIDIA Jetson Orin AGX using TensorRT‑optimized models.
arXiv:2608. 13141v1 Announce Type: cross Abstract: Vision Transformers (ViTs) demonstrate exceptional performance in computer vision but suffer from large parameter counts and quadratic computational complexity, severely limiting their deployment on resource-constrained edge hardware.
arXiv:2607. 00774v1 Announce Type: cross Abstract: Recent recursive Transformer studies have primarily reused shared parameters across computation steps to construct compact, parameter-efficient models.
arXiv:2608. 14586v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models are becoming a promising paradigm for autonomous driving, but their deployment on existing vehicle platforms remains difficult because they introduce both high inference latency and strong GPU-side resource pressure.
arXiv:2610.01905v1 Announce Type: new Abstract: Online vectorized HD map construction is essential for scaling safe autonomous driving and requires accurate, real-time inference. Prior methods typica...