Computer vision

Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.

3,375 stories · RSS feed

arXiv Computer Vision
Sep 24

DualStabSleepNet: A Dual-Domain Diffusion Stabilization Network for Robust Sleep Staging

DualStabSleepNet (DSSNet) is a dual-domain diffusion stabilization network designed to improve the robustness of automatic sleep staging across heterogeneous recording conditions. It employs a continuous-scale diffusion-based module to suppress noise while preserving physiological signals, then transforms stabilized signals into time-frequency representations for a Vision Transformer backbone. A teacher‑student guided diffusion feature stabilization further reduces feature drift, achieving state‑of‑the‑art accuracy on four public PSG datasets and demonstrating strong performance under cross‑dataset distribution shifts.

By Chongjian Wang, Chen Liu, Junjie Gao, Xiaofang Zhong, Shiyuan Han, Tong Zhang
arXiv Computer Vision
Sep 24

LiAM-SAM: Lifecycle-Aware Memory for Robust SAM2-Based MOT

LiAM‑SAM is a lifecycle‑aware memory framework designed to improve segmentation‑based multi‑object tracking (MOT) with the SAM2 foundation video model. It addresses three common failure modes—faulty track initiation, memory drift during close interactions, and unreliable re‑identification after occlusion—by introducing contrastive track initiation, motion‑ and geometry‑grounded memory correction, and adaptive context memory. The system achieves state‑of‑the‑art HOTA and IDF1 scores, with ablations showing significant gains in association metrics and a 96% reduction in identity switches.

By Gr\'egoire Francisco, Alessandro D'Amico, Samuele Costantini, Gianpiero Francesca, Lorenzo Garattoni
arXiv Computer Vision
Sep 24

VLMs Can Describe, But Not Measure: Object-Centric Scene Understanding for Robotic Manipulation

The paper introduces a modular perception framework that uses vision‑language models (VLMs) to annotate object‑level regions from a single RGB‑D observation, then grounds these annotations with depth data to build an object‑centric representation. Experiments on 151 tabletop scenes demonstrate that this decomposition maintains strong semantic performance while significantly improving localization and depth estimation compared to direct VLM inference. The resulting representation is integrated into a task‑planning system for robotic manipulation.

By Enrico Saccon, Tommaso Faraci, I\~{n}igo De La Ossa Zarzuelo, Luigi Palopoli, Marco Roveri, Matteo Saveriano
arXiv Computer Vision
Sep 24

OD3: Optimization-free Dataset Distillation for Object Detection

OD3 introduces an optimization‑free dataset distillation framework tailored for object detection. The method first iteratively places object instances in synthesized images, then screens candidates with a pre‑trained observer model to discard low‑confidence objects. Applied to MS COCO and PASCAL VOC, OD3 achieves compression ratios from 0.25% to 5% and surpasses previous detection‑focused distillation methods by over 14% on COCO mAP50 at a 1.0% compression ratio.

By Salwa K. Al Khatib, Ahmed ElHagry, Shitong Shao, Zhiqiang Shen
arXiv AI
Sep 24

NV-Reason-CT: 3D Visual Language Model for CT Analysis

NV-Reason-CT is a generative vision‑language model designed for chest and abdominal CT analysis that preserves native 3D visual encoding and incorporates radiologist‑guided reasoning. The system couples a 3D vision transformer with a language model, feeding all visual tokens and their 3D coordinates directly into language decoding to maintain volumetric spatial information. Trained on a curated corpus of about 550,000 multimodal instruction examples, the model supports abnormality classification, report generation, and interactive reasoning, achieving strong performance on CT benchmarks and reducing expert interpretation time by 50%.

By Andriy Myronenko, Dong Yang, Yucheng Tang, Baris Turkbey, Benjamin Simon, Stephanie Harmon, Rikhil Makwana, Mariam Aboian, Sena Azamat, Ibrahim Ethem Hamamci, Sezgin Er, Bjoern Menze, Marc Edgar, Yufan He, Pengfei Guo, Daguang Xu
arXiv AI
Sep 24

Do Center Biases Propagate? Robustness of Pathology Foundation Models in Whole-Slide Image Classification

The study investigates whether pathology foundation models (PFMs) carry center-related biases into whole-slide image (WSI) classification. By training models with increasing class-center correlations and evaluating six PFMs across four datasets and two MIL aggregators, the authors introduce the Area Under the Cramér's V Curve (AUCC) to measure both accuracy and degradation due to spurious correlations. Results reveal that center information propagates to WSI predictions, with robustness varying by PFM and MIL strategy, and that ComBat harmonization does not consistently improve robustness.

By Il\'an Carretero, Pablo Meseguer, Roc\'io del Amor, Valery Naranjo
arXiv AI
Sep 24

Parameter-Efficient Construction of the Rashomon Slice for Concept Bottleneck Models

The paper introduces a method for efficiently exploring the Rashomon set of Concept Bottleneck Models (CBMs) by using a parallel parameter‑efficient adaptation module, checkpointing, and a concept diversity objective. This approach generates multiple equally accurate CBMs from a single training process, achieving greater diversity than baseline methods while consuming less memory. The resulting diverse models enable trustworthy selection, reduce inter‑class confusion, and support reliable abstention in decision‑making.

By Shihan Feng, Cheng Zhang, Michael Xi, Ethan Hsu, Lesia Semenova, Chudi Zhong
arXiv AI
Sep 24

MessyKitchens: Contact-rich object-level 3D scene reconstruction

MessyKitchens introduces a new dataset of cluttered real-world kitchen scenes with detailed 3D object shapes, poses, and accurate contact information. The authors extend the SAM 3D single-object reconstruction method with a Multi-Object Decoder (MOD) to jointly reconstruct entire scenes, achieving better registration accuracy and reduced inter-object penetration compared to prior work. The dataset, benchmark, code, and pretrained models will be publicly released on the project website.

By Junaid Ahmed Ansari, Ran Ding, Fabio Pizzati, Ivan Laptev
arXiv Computer Vision
Sep 24

Tackling fluffy clouds: robust agricultural field boundary delineation from Sentinel-1 and Sentinel-2 satellite image time series

arXiv:2409.13568v3 Announce Type: replace Abstract: Accurate delineation of agricultural field boundaries is essential for effective crop monitoring and resource management. However, competing method...

By Foivos I. Diakogiannis, Zheng-Shu Zhou, Jeff Wang, Gonzalo Mata, Dave Henry, Roger Lawes, Amy Parker, Peter Caccetta, Suzanne Furby, Rodrigo Ibata, Ondrej Hlinka, Jonathan Richetti, Kathryn Batchelor, Chris Herrmann, Andrew Toovey, John Taylor