arXiv Computer Vision

FUSEye: Training-Light Fisheye Detection with Overlapping Views and Zero-Initialized Adapters

FUSEye is a lightweight training framework that adapts a frozen COCO‑pretrained YOLO26‑x detector for fisheye images by adding only about 227k parameters. It introduces three modules: GridViews to enlarge compressed boundary regions, Z‑Adapters to correct distortion‑induced feature misalignment, and AgreeFusion to fuse detections across overlapping views. On the WoodScape fisheye benchmark, FUSEye boosts YOLO26‑x mAP50 from 0.148 to 0.266 while requiring only 25% of the labeled data to achieve 0.2597 mAP50, and it also improves YOLOv8‑11 detectors.

arXiv Computer Vision
Aug 31

From Perspective to Fisheye Depth Estimation and Open-Vocabulary Segmentation

The paper introduces Distortion Extenders (DEX), learnable parameters that adapt vision foundation models to fisheye cameras by modeling distortion coefficients and correcting distributional shifts between fisheye and perspective images. DEX is applied to monocular depth estimation and open‑vocabulary segmentation across convolutional and Transformer architectures, consistently outperforming baselines on indoor and outdoor fisheye datasets. Additionally, DEX activations can be decoded to obtain distortion coefficients, aiding camera calibration.

By Rit Gangopadhyay, Alex Wong
arXiv AI
Sep 17

Ultralytics YOLO Evolution: An Overview of YOLO27, YOLO26, YOLO11, YOLOv8, and YOLOv5 Object Detectors for Computer Vision and Pattern Recognition

This paper provides a detailed overview of the Ultralytics YOLO family from YOLOv5 to YOLO27, highlighting key architectural changes, benchmarking results, and deployment considerations. It discusses the evolution of each version—YOLO27’s scale‑adaptive dual architecture, YOLO26’s loss and optimization refinements, YOLO11’s efficiency focus, YOLOv8’s anchor‑free detection, and YOLOv5’s modular ecosystem—alongside performance metrics on COCO and latency on TensorRT. The review also surveys applications in robotics, agriculture, surveillance, and manufacturing, and outlines future challenges such as dense scene handling, CNN‑Transformer integration, and hardware‑aware optimization.

By Ranjan Sapkota, Manoj Karkee
arXiv Computer Vision
Sep 11

Does YOLO26 Truly Offer Advantages Over Its Predecessors for Edge Deployment? A Benchmark Study in Aquaculture

The paper evaluates the new YOLO26 architecture, which offers NMS-free end-to-end inference and is tailored for CPU-based edge devices, against three earlier Ultralytics models (YOLOv5u, YOLOv8, and YOLO11) in aquaculture fish mortality detection. Across nano, small, and medium scales, all models achieved similar detection accuracy on a full dataset, but differences emerged in data efficiency and deployment performance: YOLOv8 reached 90% mAP50 with only 400 images, while YOLO26 variants needed 1,000 images; YOLO26n was fastest on a Raspberry Pi 5 (7.51 FPS), whereas YOLOv5mu led on CPU-based hardware. The study concludes that architectural novelty alone does not dictate suitability for edge AI in aquaculture; training data size, target hardware, and inference needs must be jointly considered.

By Rakesh Ranjan, Gajanan S. Kothawade, Kata Sharrer, Scott Tsukuda, Christopher Good
Hugging Face Trending Papers
Aug 19

SPARC: Subspace Position-Aware Robust Few-Shot Calibration for Distribution-Shifted Industrial Anomaly Detection

SPARC is a few‑shot calibration technique for vision‑based industrial anomaly detectors that corrects deployment‑time nuisances by projecting patch features onto a per‑cell subspace, requiring only up to eight verified‑normal images and no gradient updates. It operates between the encoder and detector, using a closed‑form, spatially indexed estimate based on the encoder’s native patch grid. Across seven detectors on shift‑prone benchmarks, SPARC boosts pooled Image AUROC by 13.8 pp and AU‑PRO₀.₃ by 3.5 pp, while showing modest changes on benchmarks without engineered shift.

arXiv Machine Learning
Sep 25

Albireo: Adaptive, Energy-Efficient Inference Framework for Video Object Detection on the Edge

Albireo is an adaptive, energy‑efficient inference framework for video object detection on edge devices that wraps existing detectors without modification. It uses a 10‑dimensional Kalman filter per active object to decide when to skip detector calls, predicting bounding boxes on skipped frames at near‑zero GPU cost. Evaluated on BDD100K with YOLO and RF‑DETR detectors on NVIDIA Jetson AGX Thor and Orin, Albireo maintains AP@50 within ±1.2 pp of full‑frame inference while reducing energy consumption by 12.1–17.6 % and improving accuracy for some models.

By Amir Taherin, Jos\'e Cano, Bin Ren, Yanzhi Wang, David Kaeli