Computer vision

Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.

3,375 stories · RSS feed

arXiv Computer Vision
Sep 24

ICM: Intra-class Mixing for Domain Adaptation in Adverse Weather

The paper introduces ICM, an Intra-Class Mixing Consistency framework for unsupervised domain adaptation in semantic segmentation under adverse weather. ICM mixes regions within the same image and semantic class to maintain realistic layouts, contrasting prior methods that combine across images or domains. On the Cityscapes → ACDC benchmark, ICM achieves 75.7% mIoU, surpassing previous state‑of‑the‑art results by 1.9 percentage points.

By Boying Li, Chang Liu, Britta Ayano Wilde, Gy\"orgy Kov\'acs, Tosin Adewumi, Bj\"orn Backe, Hamam Mokayed
arXiv Computer Vision
Sep 24

A generalizable structural brain MRI foundation model built through dual-priority federated pretraining

BrainFedFM is a structural brain MRI foundation model that was federatively pretrained on 164,707 3‑D scans from 42 sites using a dual‑priority approach that emphasizes informative anatomical regions locally and prioritizes site contributions globally. The model outperformed seven baseline models—including four centralized foundation models—across 20 downstream tasks (classification, regression, segmentation), achieving a mean rank of 1.68 and a 50% performance gain, especially in classification and regression and among underrepresented populations. These results demonstrate the model’s generalizability and show that federated pretraining can effectively develop neuroimaging foundation models without pooling raw images.

By Zhen Yu, Yang Liu, Xiahai Zhuang, Qingchao Chen
arXiv Computer Vision
Sep 24

CasCVS-Net: A Staged Multi-Task Cascade for Critical View of Safety Assessment

CasCVS‑Net is a staged multi‑task cascade that jointly performs object detection, semantic segmentation, and Critical View of Safety (CVS) assessment for laparoscopic cholecystectomy. The model couples tasks through predicted anatomy—boxes guide segmentation and masks provide region‑level features for CVS classification—allowing CVS assessment to rely solely on model predictions. Trained on the Endoscapes dataset, CasCVS‑Net outperforms state‑of‑the‑art methods, achieving higher mAP and mIoU scores across detection, segmentation, and CVS tasks, especially for rare hepatocystic structures.

By Bock-Zien Toh, Yuanchuan Ren, Tay Aw Yu, Ng Khee Ong, Zhehua Mao, Sophia Bano
arXiv Computer Vision
Sep 24

DMM-Align: Closed-Loop Optimization for 2D-3D Registration with Dual-Role Diffusion

DMM-Align introduces a closed‑loop framework for 2D‑3D registration that jointly refines correspondences, estimates pose, and learns representations using a shared differentiable geometric state. The method employs two diffusion processes: a geometry‑aware diffusion that improves the soft matching matrix for robust correspondence estimation, and a geometry‑conditioned diffusion teacher that feeds pose‑induced supervision back into feature learning. Experiments on 7‑Scenes and RGB‑D Scenes V2 show that DMM‑Align outperforms strong baselines, particularly in low‑overlap and heavily occluded scenarios, demonstrating the value of closed‑loop geometric feedback.

By Chongjian Wang, Junjie Gao
arXiv Computer Vision
Sep 24

MotionSpec: Spectral Trajectory Supervision for Motion-Consistent Video Generation

MotionSpec introduces a motion supervision framework for text-to-video generation that focuses on Spectral Trajectory Consistency (STC). STC builds dense anchor-relative motion trajectories, transforms them into spectral volumes, and aligns their amplitude and phase with target trajectories to constrain motion strength and temporal organization. The framework also adds Local Flow Consistency (LFC) to stabilize local motion transitions, resulting in improved motion consistency, temporal coherence, and plausibility while maintaining visual fidelity.

By Ziqi Ni, Rui Li, Shiqi Jiang, Wei Zhou
arXiv Computer Vision
Sep 24

Privacy-Preserving Semantic Segmentation from High-Resolution Depth and Ultra-Low-Resolution RGB

The paper introduces a privacy‑preserving approach for semantic segmentation that fuses high‑resolution depth with ultra‑low‑resolution RGB images. A joint 2D framework uses depth to guide RGB reconstruction and RGB‑D segmentation, while an end‑to‑end 2D‑to‑3D pipeline consolidates 2D features for 3D segmentation. Experiments on ScanNet demonstrate superior 2D and 3D performance compared to other privacy‑preserving methods, strong zero‑shot transfer to SUN RGB‑D and SceneNN, and reduced recoverability of sensitive data, with real‑robot tests showing effective object‑goal navigation.

By Xuying Huang, Swithinraj Moses Daniel, Sicong Pan, Sebastian Houben, Maren Bennewitz
arXiv Computer Vision
Sep 24

Bend the Clock: Predicting Ahead to Beat Latency in Event-Based Object Detection

The paper introduces ChronoFuse, a causal availability-time detector that predicts object states at the time its output becomes available rather than at the observation timestamp, addressing the latency mismatch in event-based multi-object detection. ChronoFuse performs lightweight cross-time fusion over a multi-scale feature hierarchy, adding only 0.17 M parameters and 0.84 ms latency overhead. It recovers a large portion of accuracy lost to latency, achieving up to 20.95 sAP on EV‑Flying data compared to 2.25 sAP for the strongest standard detector.

By Biswadeep Sen, Benoit R. Cottereau, Nicolas Cuperlier, Terence Sim
arXiv Computer Vision
Sep 24

Semantic-Guided Fusion Network for Multi-Source Remote Sensing Image Classification

The paper introduces SGFNet, a Semantic‑Guided Fusion Network for classifying multi‑source remote sensing images. It features a Semantic Mixing Convolution Block that generates semantic‑aware kernels based on contextual relationships, and a Frequency Modulated Fusion Block that fuses cross‑modal information in the frequency domain to mitigate spatial misalignment. Experiments on the Augsburg and Houston 2018 datasets show SGFNet consistently outperforms state‑of‑the‑art methods.

By Yuwei Zhao, Chuanzheng Gong, Baogui Huan, Feng Gao, Junyu Dong, Qian Du
arXiv Computer Vision
Sep 24

Automated Palynological Analysis System: Integrating Deep Metric Learning, Detection and Classification in Bright Field Microscopy

The paper introduces an automated high‑throughput microscopy system for melissopalynology that combines H∞ robust mechanical control with deep learning pipelines. It uses U^2‑Net for salient object detection and a DINOv2 Vision Transformer trained via deep metric learning for pollen grain classification, augmented with Gradient‑Weighted Attention for interpretable texture annotations. The system reports a 95.8% classification recall and at least a six‑fold speedup over manual expert analysis.

By J. Staforelli-Vivanco, R. Jofr\'e, P. Coelho, I. Sanhueza, L. Viafora, C. Toro, J. Troncoso, M. Rondanelli-Reyes, I. Lamas, Andy Banegas-Medina, Isis-Yelena Montes, B. Mu\~noz-Cepeda, V. Salamanca-Levi, M. Gonz\'alez-Ortiz, E. Vera
arXiv Computer Vision
Sep 24

CRISP: Compositional Relations as Invariant Structural Priors for Domain Generalization

CRISP (Compositional Relational Invariance from Spatial Primitives) is an image‑classification framework that decomposes visual recognition into primitive elements and their relational composition. It represents these compositions with soft unary, binary, and ternary predicates over primitive locations and appearance, enabling differentiable spatial and visual alignment learned end‑to‑end. Evaluated on five DomainBed datasets covering style, provenance, and camera‑trap shifts, CRISP achieves new state‑of‑the‑art performance on both benchmarks.

By Dat Nguyen, Duc-Duy Nguyen
arXiv Computer Vision
Sep 24

Know-Your-Scene (KYS)-SLAM: Hierarchical Semantic-Motion Priors for Feature Matching in Stereo Visual SLAM

Know-Your-Scene (KYS)-SLAM extends ORB‑SLAM3 by replacing binary feature rejection with continuous correspondence modulation based on semantic, panoptic, and motion priors. Each keypoint is augmented with hierarchical compatibility scores that down‑weight features on independently moving objects while preserving static structure, using a training‑free depth‑aware ego‑motion model and self‑calibrating thresholds. Across 21 stereo sequences, KYS‑SLAM achieves a 17.4% ATE RMSE reduction on outdoor KITTI, 27.7% on indoor EuRoC, and significant improvements on dynamic and synthetic datasets without per‑sequence tuning.

By Preeti Chatterjee, Jin Lu, Jin Sun, Suchendra M. Bhandarkar
arXiv AI
Sep 24

Learning Spectral Allocation: A Fractional Diffusion Framework for Adaptive Volumetric Segmentation

The paper introduces FHEAT, a fractional diffusion operator derived from the discrete cosine transform, to learn how much spectral mixing each stage of a 3D medical segmentation network should perform. By reparameterizing the operator with a semigroup time, the authors enable the optimizer to decide whether global mixing is needed, resulting in a lightweight U‑shaped architecture (Light‑UNETR) paired with a Kolmogorov‑Arnold mixer (KAN3D). In semi‑supervised experiments, FHEAT‑Seg achieves state‑of‑the‑art Dice scores while dramatically reducing FLOPs through learned spectral sparsification.

By Yi-Hui Shen, Tie-Qiang Li
arXiv AI
Sep 24

Anon: Extrapolating Adaptivity Beyond SGD and Adam

The paper introduces Anon, an optimizer that extends adaptivity beyond the traditional bounds of SGD and Adam by allowing extrapolation across the entire real-number spectrum. It addresses the limitations of prior tunable optimizers that only interpolate between 0 and 1 adaptivity, showing that optimal adaptivity can require negative values for CNNs or values greater than one for Transformers. Anon incorporates Incremental Delay Update (IDU) to maintain provable stability and demonstrates competitive performance on image classification, diffusion, and large language modeling tasks.

By Yiheng Zhang, Kaiyan Zhao, Shaowu Wu, Yiming Wang, Jiajun Wu, Leong Hou U, Steve Drew, Xiaoguang Niu
arXiv AI
Sep 24

Evaluating ADC-only deep learning pipelines for breast cancer detection and segmentation using standalone diffusion-weighted MRI

The paper evaluates deep learning pipelines that use only apparent diffusion coefficient (ADC) maps from diffusion-weighted MRI (DW-MRI) for breast cancer detection and segmentation. It compares various state‑of‑the‑art models and claims to be the first comprehensive assessment of ADC‑only approaches for classification, detection, and segmentation tasks. The study highlights the potential of DW-MRI as a faster, contrast‑free alternative to dynamic contrast‑enhanced MRI for breast cancer imaging.

By Pablo Garc\'ia Marcos, Paula Puerta Gonz\'lez, Guillermo Lorenzo, H\'ector G\'omez, Covadonga del Camino, Angel Rio-Alvarez, V\'ictor M. Gonz\'alez
arXiv Machine Learning
Sep 24

A Scaling Study for fMRI Foundation Models

The study investigates how data volume, model size, and training duration affect the performance of fMRI foundation models. Using over 200 datasets and 10,000 GPU‑hours, the authors find that larger models benefit more from additional data, and that at a fixed compute budget, increasing data yields greater gains than enlarging the model. By selecting optimal combinations of data, size, and duration, they produce models that outperform existing fMRI foundation models on out‑of‑distribution tasks while requiring less pretraining compute.

By Wenhao Ye, Xuanye Pan, Junfeng Xia, Junxiang Zhang, Mo Wang, Quanying Liu
arXiv Machine Learning
Sep 24

Integrated Multivariate Segmentation Tree for Heterogeneous Credit Data Analysis in Small- and Medium-Sized Enterprises

The paper introduces the Integrated Multivariate Segmentation Tree (IMST), a new framework that combines financial data and textual information for credit evaluation of small- and medium-sized enterprises. IMST transforms text into numerical matrices via matrix factorization, selects key financial features with Lasso regression, and builds a multivariate segmentation tree using Gini or entropy with weakest-link pruning. Experiments on 1,428 Chinese SMEs show an 88.9% accuracy, outperforming baseline decision trees, SVMs, and neural networks while offering better interpretability and computational efficiency.

By Lu Han, Xiuying Wang
arXiv Computer Vision
Sep 24

S2A:Semantic-to-Spatial Alignment for Alignment-Free RGB-T Salient Object Detection

The paper introduces S2A, a semantic-to-spatial alignment framework designed for alignment‑free RGB‑T salient object detection. It employs a global‑guided hierarchical fusion module to refine intra‑modal features, an alignment‑free cross‑modal channel attention module to exchange semantic information, and a spatial deformable cross‑attention module to recover local spatial correspondence. These components collectively reduce misalignment‑induced feature contamination and achieve competitive performance on public benchmarks without additional bells and whistles.

By Qiangqiang Zhou, Yang Luo, Yong Chen, Jiawei Xu
arXiv Computer Vision
Sep 24

Overlapping Visual Grouping Without Semantic Priors

The paper introduces Domain Parent Grouping (DPG), a sensor‑grounded method that forms perceptual units directly from raw measurements without relying on semantic priors. DPG operates across three domains—local luminance, direct chromatic, and contextual chromatic—creating spatially connected groups that overlap across domains, thus producing a non‑exclusive grouping representation. Experiments on the BSDS500 dataset show that DPG’s groups align with low‑level image structure and correlate with human‑annotated regions and boundaries.

By Teemu Saukkio, Hashem Haghbayan, Juha Plosila
arXiv Computer Vision
Sep 24

Hybrid Gaussians for Robust Open-Vocabulary 3D Segmentation with Multi-View Object Association and Boundary Refinement

Hybrid Gaussians is a new 3D representation that jointly models object association and language-aligned semantics for open‑vocabulary 3D segmentation. It uses a Multi‑View Object Association mechanism that combines Observation Fusion and Semantic Contrastive Learning to improve identity consistency and semantic discrimination, and a Boundary Reconstruction Optimization to refine local boundary structure. Experiments on LERF and 3D‑OVS show strong quantitative and qualitative results, achieving 59.1% mIoU on LERF, a 13.4% relative gain over the baseline.

By Xueqi Qiu, Yueming Sun, Tianyu Zhang, Yuxuan Xia, Yang Long
arXiv Computer Vision
Sep 24

SGDet3D++: Geometry-Grounded Semantics for 4D Radar and Camera 3D Object Detection

SGDet3D++ introduces a geometry‑grounded approach to 4D radar‑camera 3D object detection by explicitly conditioning evidence on evolving object hypotheses. It employs Anchor‑Grounded Semantic Retrieval, Geometry‑Consistent Anchor Refinement, and Doppler‑Verified Correspondence to filter and align semantic, geometric, and temporal cues before updating queries. The method achieves significant performance gains on OmniHD‑Scenes, ManTruckScenes, and TJ4DRadSet, with detailed ablations showing improvements in occlusion handling, target‑return purity, and motion consistency.

By Xiaokai Bai, Zhenyu Fan, Lianqing Zheng, Songkai Wang, Si-Yuan Cao, Hui-liang Shen