Computer vision

Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.

3,374 stories · RSS feed

arXiv AI
3d ago

Supermarket Product Detection and Recognition: Utilizing Deep Learning with Rectified Imagery

The paper investigates how rectifying supermarket product images using homography estimation and the Hough transform can improve deep learning-based object detection. It evaluates the impact of angle variation and object density on detection accuracy, highlighting both benefits and limitations of image rectification. The authors advocate for a new dataset to further study these effects.

By Mayank Sah, Jimson Mathew
arXiv Computer Vision
3d ago

Con-DSO: Learning Short-Horizon Consistency Priors for RGB-D Direct Sparse Odometry

Con-DSO introduces a consistency-aware RGB‑D direct sparse odometry framework that learns pixel‑level photometric and geometric uncertainty from adjacent RGB‑D frame pairs. The network predicts uncertainties that are converted into pairwise quality scores, guiding support‑pixel selection and forming a host‑side quality prior for keyframe tracking. Experiments on five public benchmarks show that this approach reduces absolute trajectory error by over 20% on ICL‑NUIM and by 50–80% on other datasets, improving robustness in challenging environments.

By Haolan Zhang, Thanh Nguyen Canh, Chenghao Li, Ziyan Gao, Xiongwen Jiang, Nak Young Chong
arXiv Machine Learning
3d ago

Skillful Data-Driven Subseasonal Soil Moisture Forecasting: Prospects and Limits for Flash Drought Prediction

The paper presents a Vision Transformer-based model for subseasonal soil‑moisture forecasting over Europe, showing that forecast skill depends heavily on how the prediction problem is formulated. By using residual learning and forecasting root‑zone soil moisture in physical units, the model outperforms persistence and existing deep‑learning and ECMWF baselines, providing well‑calibrated probabilistic predictions. However, predicting flash drought onset—defined by rapid multi‑pentad intensification—remains a challenge shared by all current subseasonal‑to‑seasonal systems.

By Noelia Otero, Atahan \"Ozer, Miguel-\'Angel Fern\'andez-Torres, Jackie Ma
arXiv Computer Vision
3d ago

WiSPER: Pose-Supervised Predictive and Residual Flow Refinement For Multi-Person 3D Pose Estimation With WiFi CSI

WiSPER is a two‑stage framework for multi‑person 3D pose estimation using WiFi channel state information (CSI). The first stage, Pose‑Aware Masked Embedding Learning (PAMEL), couples masked latent prediction with pose‑set supervision to guide the encoder toward joint localization from partial observations. The second stage, Residual Flow refinement with Transformer (ReFT), generates pose candidates for a variable number of people and refines each candidate through a conditional flow guided by coarse coordinates and decoder features. Trained with paired CSI and pose annotations, WiSPER achieves a mean per‑joint position error of 63.72 mm on the PiW3D dataset, a 40.0 % improvement over WiFi‑JEPA and significant reductions for two‑ and three‑person scenarios.

By Gabriel Lee Jun Rong, Shanhong Liu, Pai Chet Ng, Konstantinos N. Plataniotis, Jamal Seyedmohammadi, S. Mohammad Sheikholeslami
arXiv Computer Vision
3d ago

Towards benchmarking Western Bluebird detection in the wild

The paper introduces a new benchmark dataset of over 6,000 high‑resolution images for detecting and segmenting Western bluebirds in natural, cluttered scenes. Experiments show that supervised detectors such as Faster R‑CNN and RT‑DETR outperform open‑vocabulary models in zero‑shot settings, though fine‑tuned YOLO‑World can compete. Segmentation results favor Mask R‑CNN for mask quality, while YOLOv8‑Seg offers the best precision and speed. The study also identifies multiple factors—scale, brightness, contrast, clutter, blur, crowding, and session variation—as contributors to detection failures.

By Estela Monserrat Arriaga Santana, Julian Rosas Scull, Ibeth P. Alarc\'on, Bibiana Montoya, Aylin Sosa Mej\'ia, Hugo Jair Escalante
arXiv Computer Vision
3d ago

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs

The paper introduces FindIt, the first comprehensive benchmark for evaluating the promptable localization abilities of generalist multimodal large language models (MLLMs). It covers four core task categories—object detection, referring expression detection, instance-level detection, and video-based detection—and provides a unified framework that standardizes inputs, enforces parsable bounding box outputs, and defines transparent evaluation protocols. Using this benchmark, the authors assess a range of open-source and proprietary MLLMs, revealing both their strengths and limitations, particularly their sensitivity to formatting constraints and difficulty generalizing to minor variations.

By Eshika Khandelwal, Jingjing Pan, Mingfang Zhang, Quan Kong, Lorenzo Garattoni, Hilde Kuehne
arXiv Computer Vision
3d ago

Toward Open-World Video Segmentation over Long Horizons

The paper introduces Savvy, a zero‑shot, semi‑online, class‑agnostic system that persistently discovers objects and maintains their identities in long videos, and OGA, an evaluation suite that rewards coherent part‑level predictions even when their granularity differs from reference annotations. Savvy combines modular mask discovery, deferred admission based on accumulated evidence, and track consolidation to keep an evolving object set, outperforming DEVA+SAM and EntitySAM on ScanNet and HM3D datasets in metrics such as VPQ_inf, STQ, and AQ. OGA further distinguishes coherent part‑level support from temporal identity failures, revealing that conventional one‑to‑one VPQ metrics are sensitive to annotation granularity and that temporal failures can be detected even when frame‑level masks remain unchanged.

By Qing Su, Kaiyang Li, Yuan Zhuang, Fei Miao, Shihao Ji
arXiv Computer Vision
3d ago

Multitask Conditional Generative Adversarial Network Enables Automatic Whole Knee Cartilage and Menisci Segmentation and Reliable $T_{1\rho}$ and $T_2$ Quantification Without High-Resolution Morphological Images

The study introduces a multitask conditional generative adversarial network (MT‑cGAN) that simultaneously synthesizes high‑resolution DESS‑like images and segments knee cartilage and menisci directly from quantitative MRI echo images. Evaluated on 508 knee MRIs from 361 subjects, MT‑cGAN achieved a mean Dice score of 0.84 for segmentation and the lowest coefficient of variation for T1ρ (1.84%) and T2 (1.81%) quantification, outperforming existing conditional GAN approaches. By eliminating the need for separate high‑resolution morphological scans, the method shortens scan times and supports clinical adoption of quantitative MRI.

By Ahmed Tahseen Minhaz, Richard Lartey, Zhiyuan Zhang, Jeehun Kim, Kunio Nakamura, Mingrui Yang, Jiasen Zhang, Weihong Guo, Naveen Subhas, Carl S. Winalski, Xiaojuan Li
arXiv Machine Learning
3d ago

Uncertainty Quantification Is Indispensable for Reliable Connectome-Based Graph Learning: A Narrative Review and Case Study

The paper reviews uncertainty quantification (UQ) methods for graph neural networks used in connectome-based diagnostic classification and presents a case study on a temporal Graph Attention Network applied to dynamic functional connectivity data for Cocaine Use Disorder. It highlights that deterministic GNNs can produce overconfident predictions, as shown by a Monte Carlo dropout audit revealing high confidence on misclassified subjects. The study demonstrates the need for rigorous UQ, calibration, and selective prediction to ensure reliable graph-based biomarkers in clinical neuroscience.

By Mansooreh Pakravan
arXiv Machine Learning
3d ago

PRUE: A Practical Recipe for Field Boundary Segmentation at Scale

arXiv:2603.27101v2 Announce Type: replace-cross Abstract: Large-scale maps of field boundaries are essential for agricultural monitoring tasks. Existing deep learning approaches for satellite-based f...

By Gedeon Muhawenayo, Caleb Robinson, Subash Khanal, Zhanpei Fang, Isaac Corley, Alexander Wollam, Tianyi Gao, Leonard Strnad, Ryan Avery, Lyndon Estes, Ana M. T\'arano, Nathan Jacobs, Hannah Kerner