Computer vision

Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.

3,374 stories · RSS feed

arXiv AI
Oct 1

Raw Imagery Impacting Your AI: Should You Care?

The paper investigates how raw or minimally processed satellite imagery affects onboard AI object detection for space missions. By systematically degrading Very High Resolution Maxar images in terms of Signal‑to‑Noise Ratio, Modulation Transfer Function, and Ground Sampling Distance, the authors evaluate three lightweight detectors—YOLOv5s, YOLOX‑S, and NanoDet—on the resulting data. Results show that image quality impacts detection performance in a degradation‑specific way, with GSD consistently shifting performance, while MTF and SNR effects vary by model and resolution; severe blur‑plus‑noise combinations cause the greatest losses.

By Adrien Dorise, Marjorie Bellizzi, St\'ephane May
arXiv Computer Vision
Oct 1

Consensus-Aware Multi-Source Fusion for Reference-Guided Camouflaged Object Detection

The paper introduces a consensus‑aware multi‑source fusion framework for reference‑guided camouflaged object detection. It couples trainable PVTv2 query features with frozen DINOv3 representations, using reference‑conditioned correlation to select foundation‑model evidence before multi‑scale fusion. The method also aggregates multiple references via cross‑reference consensus aggregation and injects reference information at semantic depths matched to the query features, achieving complementary improvements in experiments.

By Junyang Xia, Luocheng Zhang, Wenwen Pan, Chifeng Zhu, Yang Yang, Xinchun Liu, Jiajun Ding
arXiv Computer Vision
Oct 1

Event-Driven Refresh and Recurrence Memory to Reduce Stale Grounding in Referring Video Object Segmentation

The paper introduces Event-Driven Refresh + Recurrence Memory (EDRRM) to improve Referring Video Object Segmentation (RVOS). EDRRM selectively re-invokes the Sa2VA model at stable change points, using an EMA‑smoothed event score from tracking cues and a recurrence memory that retrieves anchor frames via CLIP similarity. Experiments on Ref‑DAVIS17, MeViS, and ReVOS show that EDRRM maintains or surpasses J&F scores while reducing refresh calls and false‑positive failures, with modest overhead compared to Sa2VA inference.

By Abu Hanif Muhammad Syarubany, Jaehyun Jang, Siwoo Lim, Seungyeon Ryu, Chang D. Yoo
arXiv Computer Vision
Oct 1

PCB-MC: Missing Component Analysis in Printed Circuit Boards

PCB-MC is a new dataset for detecting missing components on printed circuit boards, featuring 197 distinct board types with footprint-level annotations derived from the RF100 dataset. The dataset includes multiple augmented samples per board type and introduces board type-aware cross‑validation splits to prevent layout leakage between training and test sets. Benchmark results show that supervised models suffer high false‑negative rates on unseen board designs, while unsupervised anomaly detection methods fail due to misalignment with board-specific references, highlighting the remaining challenges in missing component detection across diverse PCB layouts.

By Betsy Villa Brochero, Ian Gibson, Estefania Talavera
arXiv AI
Oct 1

Learning Continuous Neural Representation of Stochastic Hybrid Systems

The paper demonstrates that a stochastic hybrid system (SHS), which combines continuous dynamics governed by a stochastic differential equation (SDE) with discrete resets triggered by a Markov kernel, can be approximated by a single SDE in a higher‑dimensional latent space. By encoding the reset branches with auxiliary variables, the resets become deterministic, allowing the system’s manifold to be glued and embedded into Euclidean space. This embedding eliminates the need for explicit reset terms in the hybrid Fokker‑Planck equation, and the authors propose a loss function that matches evolving state distributions, enabling the latent SDE to recover the SHS’s probability evolution without mode labeling, trajectory segmentation, or event‑based simulations.

By Sangli Teng, Hang Liu, Koushil Sreenath
arXiv AI
Oct 1

Colorectal Cancer Segmentation with Adaptive Augmentation and Multi-Resolution Ensemble Models

The paper presents an automated segmentation pipeline for whole‑slide histopathology images of colorectal cancer, labeling tumor grades 1‑3 and normal mucosa. It employs dense prediction transformers with multiple encoder backbones, overlapping patches, test‑time augmentation, and an adaptive augmentation policy guided by large language models. The approach, combined with soft‑voting ensembles and post‑processing refinements, raises the F1 score from 62.92 to 69.84 on a colorectal cancer grade dataset.

By \"Umit Mert \c{C}a\u{g}lar, Alptekin Temizel
arXiv AI
Oct 1

HiWE: Hierarchical World Knowledge Model with Visual Keypoint Enhancement for Zero-Shot 3D Path Planning

HiWE is a hierarchical world knowledge model that enables zero‑shot 3D path planning by linking visual grounding with language‑based planning through a point‑based interface. It uses PointVLM to map task‑relevant objects to image coordinates, lifts these predictions into a semantic 3D representation with depth data, and then a language planner (3DLLM) generates end‑effector waypoints and gripper commands. The system is evaluated on 14 simulated manipulation tasks and four physical‑robot tasks, with ablations on visual training data, spatial inputs, and grasp selection.

By Guoqing Ma, Mingqi Yuan, Chen Gao, Jiayu Chen, Shan Yu
arXiv AI
Oct 1

WinoTS: Wavelet-based Self-Distillation for Time Series Models

WinoTS introduces a wavelet‑based self‑distillation framework for time‑series models that uses time‑frequency augmentations to create multi‑scale structural views, avoiding distortion of signal dynamics. The method outperforms state‑of‑the‑art baselines in long‑term forecasting, cross‑domain zero‑shot transfer, and unsupervised anomaly detection, and linear probing on frozen representations often beats fully supervised training from scratch. Ablation studies show WinoTS is architecture‑agnostic and demonstrates that time‑frequency transformations offer a principled alternative to vision‑style spatial augmentations.

By Noam Major, Kathy Razmadze, Yoli Shavit
arXiv AI
Oct 1

GroundingPI: A Grounding Foundation Model towards Physical Intelligence with Visual Primitives

GroundingPI is a 4‑billion‑parameter grounding foundation model that generates points and boxes as quantized coordinates in a shared vocabulary. It is trained with multimodal and spatial pretraining, supervised fine‑tuning, and reinforcement learning, achieving a new state‑of‑the‑art average of 73.68% across 34 grounding benchmarks. As a visual backbone, GroundingPI improves performance in robotic manipulation and autonomous driving, outperforming larger models and mainstream backbones in several out‑of‑distribution settings.

By Qize Yu, Lianrui Fan, Boyu Chen, Jiaqi Liang, Xini Ding, Yue Chen, Zetian Song, Yuran Wang, Yi Zou, Kaixuan Wang, Tianxing Chen, Wenxuan Song, Bohan Zhou, Mingleyang Li, Siqiao Huang, Yuqi Ye, Caigao Jiang, Wei Wei, Ruihai Wu, Hang Zhang, Yixiao Ge, Shuchang Zhou, Shilong Liu, Xianming Liu, Ping Luo, Shiyu Huang