Computer vision

Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.

3,375 stories · RSS feed

arXiv Machine Learning
Sep 25

Let Training Guide Selection: Online Synthetic Data Filtering via Real-Anchored Utility

The paper introduces FROST, an online framework that filters synthetic training data by estimating its utility through gradient feedback anchored in real data. FROST calibrates batch utility against recent history to decide when to filter, removing 20–30% of synthetic samples while improving performance on image classification and LLM fine-tuning tasks. The method is also applied to a large‑scale industrial ads re‑ranking system, yielding significant gains over an optimized production baseline.

By Yanran Wu, Sana Lakdawala, Renzo Tassara Miller, Chongyang Bai, Sharath Ciddu, Shivendra Pratap Singh, Kungang Li, Sandeep Pandey, Chunwei Liu
arXiv Machine Learning
Sep 25

Multimodal Routing and Region Refinement for Language-Guided Medical Image Segmentation

Multimodal Routing and Region Refinement for Language-Guided Medical Image Segmentation (MRSeg) is a parameter‑efficient framework that uses frozen ConvNeXt‑Tiny and PubMedBERT encoders to extract multiscale visual features and clinical text tokens. A joint router predicts a sparse mixture over low‑rank adapter bases, enabling separate adaptation for two visual scales and text while keeping feature‑specific parameters distinct. Region Bridge aggregates dense visual tokens into latent regions using text‑derived queries, refines them via self‑attention and text cross‑attention, and redistributes the refined information back to the feature maps, culminating in a multiscale decoder that combines refined semantic features with shallow image evidence. MRSeg achieves state‑of‑the‑art Dice/mIoU scores on QaTa‑COV19 and MosMedData+ with only 7.11 M trainable parameters and 7.60 GFLOPs.

By Md Maklachur Rahman, Md Hasan Al Banna, Saraf Anjum, Assame Arnob, Tracy Hammond
arXiv Machine Learning
Sep 25

MLPerf Automotive

MLPerf Automotive is the first standardized public performance benchmark for evaluating machine learning systems used in automotive AI acceleration. Developed by MLCommons, it addresses the unique constraints of automotive workloads—such as sensor suites, safety, and real‑time processing—that existing benchmarks cannot handle. The benchmark covers perception tasks (2D/3D object detection, semantic segmentation), end‑to‑end driving, and infotainment, and includes curated models, methodology, and reference implementations available on GitHub.

By Radoyeh Shojaei, Predrag Djurdjevic, Mostafa El-Khamy, James Goel, Kasper Mecklenburg, P{\i}nar Muyan-\"Oz\c{c}elik, John Owens, Tom St. John, Jinho Suh, Arjun Suresh
arXiv Computer Vision
Sep 25

Token Clustering and Semantic Sequence Mamba for Hyperspectral Image Classification

Token Clustering and Semantic Sequence Mamba (STMamba) is a new approach for hyperspectral image classification that organizes sparse tokens into semantically coherent sequences. It uses a hierarchical encoder-decoder with a Token Clustering Module (TCM) to select semantic tokens and a Cross-scale Neighborhood Attention (CNA) Upsampler to restore dense features. At the micro level, density-aware clustering and a quadtree-based dynamic selection keep sparse, spatially distributed tokens, while Spatial and Spectral Semantic-wise Sequencing Mamba (SWSM) modules capture long-range spatial and spectral dependencies within homogeneous semantic token sequences. Experiments on three large-scale benchmark datasets show that STMamba outperforms state‑of‑the‑art methods in both quantitative and qualitative metrics.

By Yimin Zhu, Mahmood Elahi, Lincoln Linlin Xu
arXiv Computer Vision
Sep 25

PePESeg3D: Perception Prior Enhances Multi-Scale Segmentation for 3D Gaussian Splatting

PePESeg3D introduces perception priors into a multi‑scale 3D Gaussian segmentation pipeline, integrating monocular depth and mask constraints during geometry reconstruction and dense depth‑color cues with view‑consistent centroid supervision during contrastive feature learning. This dual‑stage approach aligns geometry with semantic structure and compensates for incomplete mask supervision from 2D foundation models. Experiments on SPIn‑NeRF, LERF‑Mask, and NVOS benchmarks show state‑of‑the‑art performance in both multi‑scale segmentation and scene reconstruction.

By Sungjae Choi, Seunghee Koh, Junmo Kim
arXiv Computer Vision
Sep 25

MEVL-STP: Multi-Encoder and Vision Language Model for Arbitrarily Shaped Scene Text Spotting

MEVL-STP introduces a two‑stage pipeline for spotting arbitrarily shaped scene text. The detection stage fuses features from six frozen vision encoders via a hierarchical Feature Pyramid Network and a Progressive Scale Expansion network to produce precise polygon masks. The recognition stage then crops these masks and feeds them to a fine‑tuned Qwen3‑VL‑8B‑Instruct model, achieving state‑of‑the‑art detection and end‑to‑end performance on CTW1500, Total‑Text, and ICDAR 2015 without synthetic pretraining.

By Aman Anand, Partha Pratim Roy, Shivakumara Palaiahnakote
arXiv Computer Vision
Sep 25

Exploiting Target Knowledge from MLLMs for Robust Few-Shot Segmentation

The paper introduces MK‑FSS, a few‑shot segmentation framework that leverages Multimodal Large Language Models (MLLMs) to extract spatial and semantic target knowledge from query images. Spatial knowledge is encoded into a memory representation and fused with support‑guided memory via a dual‑memory debate‑fusion module, while semantic knowledge is turned into a textual feature and combined with multi‑scale query features through a progressive cross‑modal prompt generator. Together, these components produce a robust target representation that improves segmentation performance over existing methods.

By Yijun Hu, Heng Fan, Libo Zhang
arXiv Computer Vision
Sep 25

MoVISA: Multi-Token Reasoning for Video Object Segmentation

MoVISA introduces Multi-Token Reasoning for Video Object Segmentation, using multiple segmentation tokens (e.g., SEG0, SEG1) instead of a single token to better localize multiple objects over time. This approach enhances fine-grained alignment between language prompts and spatio-temporal mask predictions, leading to improved performance and interpretability. On benchmarks such as MeViS, DAVIS17, ReVOS, and Ref-Youtube-VOS, MoVISA achieves significant gains, notably a 13.2% J and F improvement on MeViS and an 8.4% J and F improvement on ReVOS.

By Ruining Zhao, Ho Kei Cheng, Alexander G Schwing
arXiv Computer Vision
Sep 25

Seeing Is Not Measuring: Tool-Augmented Metric Spatial Reasoning for Vision-Language Models

The paper introduces a tool‑augmented framework that enhances a small Vision‑Language Model (Qwen3.5‑4B) with geometric tools—3D object detection, metric depth estimation, and deterministic solvers for distance, size, and bearing—to improve metric spatial reasoning. By moving metric computation from the model’s weights into explicit solvers, the approach achieves significant gains on ReVSI‑Bench tasks, notably increasing absolute distance accuracy from 0.46 to 0.74 MRA and relative direction accuracy from 25.9% to 73.4%. The modular design allows swapping in different detectors, enabling a clear separation between perception and reasoning errors, and the model can autonomously sequence the tools to match a scripted pipeline on most tasks.

By Kai Glantz, Clemens Grange
arXiv Computer Vision
Sep 25

Shadow Reduction in Ultrasound Imaging Using Differentiable Simulation and Radiance Field Decomposition

RFlash is a physics‑informed post‑processing technique that removes acoustic shadows in ultrasound by decomposing beamformed images into attenuation and scatter‑intensity maps using a differentiable radiance‑field model. It re‑renders images to eliminate depth‑dependent signal loss, effectively simulating a virtual transducer advance. Across thousands of fetal brain, abdominal, and liver scans, RFlash outperforms traditional attenuation correction, reduces prediction error by 40% in fetal brain imaging, and provides shadow‑confidence maps that enhance bone‑shadow segmentation.

By Valentin Bacher (Oxford Machine Learning in NeuroImaging Lab, University of Oxford, United Kingdom), Pak Hei Yeung (Oxford Machine Learning in NeuroImaging Lab, University of Oxford, United Kingdom, Quantitative Healthcare Analysis), Bernhard Kainz (Friedrich-Alexander-Universit\"at Erlangen-N\"urnberg, Germany, Imperial College London, United Kingdom), Madeleine K. Wyburd (Oxford Machine Learning in NeuroImaging Lab, University of Oxford, United Kingdom, Department of Computer Science, University of Copenhagen, Denmark), Nicola K. Dinsdale (Oxford Machine Learning in NeuroImaging Lab, University of Oxford, United Kingdom), Michael Gray (Institute of Biomedical Engineering, University of Oxford, United Kingdom), Ana I. L. Namburete (Oxford Machine Learning in NeuroImaging Lab, University of Oxford, United Kingdom)
arXiv Computer Vision
Sep 25

SpectralCTGaussians: Projection-Domain Reconstruction and Basis Material Decomposition for Spectral CT using 3D Gaussian Splatting

SpectralCTGaussians introduces a novel spectral CT reconstruction and basis material decomposition technique that employs 3D Gaussian splatting with per‑Gaussian material fractions and energy‑dependent basis functions. By jointly optimizing these parameters across all energy channels via a differentiable polychromatic forward model, the method achieves superior novel view synthesis and higher PSNR for spectral CT volume reconstruction compared to traditional and learning‑based baselines. It also provides one‑step material decomposition with direct RGB segmentation and recovers the photoelectric basis more accurately than existing pipelines.

By Reinout Vos, Saptarshi Neil Sinha, Michael Weinmann
arXiv Computer Vision
Sep 25

TopoFuse: Topology-Aware Tri-Planar Fusion for 3D Cryo-Electron Tomography Segmentation

TopoFuse introduces a topology-aware tri-planar fusion method for 3D cryo-electron tomography segmentation. It replaces traditional loss penalties with a differentiable projection operator that identifies and sparsely edits critical voxels to enforce specified topological constraints. The approach achieves a 54% reduction in Betti number error, a 4.6-point Dice improvement, and edits only 3.1% of voxels across three benchmarks.

By Rohit Kumar Salla, Neelesh Gupta, Xingjian Li, Min Xu
arXiv Computer Vision
Sep 25

Mind the Gap: Mesh-Guided Repair of Broken Vessels

The paper introduces a mesh-guided post‑processing framework that repairs broken vessel segmentations produced by nnU‑Net. By fitting a deformable template mesh to each predicted binary mask, the method conservatively reconnects disconnected components using thin bridge candidates while respecting foreground‑growth constraints. Evaluation on aorta, Circle of Willis, and pulmonary artery datasets shows that connectivity metrics (ccDice and Betti‑0) improve dramatically with negligible change in overall overlap (Dice).

By Gniewosz Drwiega, Wojciech Szymanski, Marek Wodzinski
arXiv Computer Vision
Sep 25

Can Frozen Hyperspherical Features Guide the Selection of Pseudo Masks?

The paper introduces SphereTrust, a method that uses frozen self‑supervised hyperspherical features to evaluate and rank candidate masks produced by foundation segmenters like SAM. By measuring angular contrast, foreground coverage, and image‑frame contact, SphereTrust can select high‑quality masks in 0.55 s per image and outperforms existing baselines on multiple segmentation tasks. The selected masks are then used as priors to train student models, improving performance on several benchmark datasets.

By Xinge Guo, Fengyang Xiao, Dingming Zhang, Yuhan Chen, Rihan Zhang, Xingjian Li, Tianyang Wang, Chunming He, Sina Farsiu
arXiv Computer Vision
Sep 25

MDE-VIO: Enhancing Visual-Inertial Odometry Using Learned Depth Priors

MDE-VIO integrates learned depth priors into the VINS-Mono optimization backend to improve visual‑inertial odometry in low‑texture environments. The framework enforces affine‑invariant depth consistency and pairwise ordinal constraints while filtering unstable artifacts with variance‑based gating, keeping computation within edge‑device limits. Experiments on TartanGround and M3ED datasets show the method prevents divergence and reduces Absolute Trajectory Error by up to 28.3%.

By Arda Alniak, Sinan Kalkan, Mustafa Mert Ankarali, Afsar Saranli, Abdullah Aydin Alatan
arXiv Computer Vision
Sep 25

COMPASS: Fusion-Matched Supervision for Missing-Modality Human Sensing

COMPASS is a completion-and-fusion framework designed for multimodal human activity recognition (HAR) and human pose estimation (HPE) when some modalities are missing at inference. It assigns each modality to a fixed slot, filling it with either an observed representation or a completion inferred from available inputs, and uses fusion‑matched supervision to train completions against real targets at the readout level. Experiments on XRF55 and MM‑Fi datasets show that COMPASS outperforms strong baselines and alternative matching strategies for both HAR and HPE.

By Hao Wang, Yanyu Qian, Pengcheng Weng, Zixuan Xia, William Dan, Yangxin Xu, Fei Wang
arXiv Computer Vision
Sep 25

ConPro: Contrast Projection Pretraining for Label-Efficient Vessel Segmentation in DSA Sequences

ConPro introduces a self‑supervised pretraining method for vessel segmentation in digital subtraction angiography (DSA) by using a contrast projection target— the normalized drop of each pixel below its temporal median. On the DIAS and DSCA datasets, ConPro outperforms training from scratch across 10%, 20%, and 50% labeled data, and it is the best among compared methods on DSCA at 20% and 50% labels. When combined with semi‑supervised training, ConPro‑derived weights boost the UniMatch baseline by 0.5–2.0 Dice points and 0.9–2.3 clDice points, achieving 75.4 Dice on DIAS and 81.3 on DSCA.

By Xinge Guo, Yuanhao Wang, Liqi Shu, Yang Liu, Min Xu
arXiv Computer Vision
Sep 25

Smartphone-Based Method for Automated Speed Enforcement

The paper presents a smartphone-based system that uses computer vision to automatically estimate vehicle speed and identify vehicles by license plate, make/model, and color. Experiments on a Brazilian dataset and real-world recordings in Austin, Texas show moderate recognition rates: 46% for license plates, 60.8% for color, 48.6% for make, and 16.89% for make/model. The study also discusses legal, technological, and practical considerations for deploying such smartphone recordings in traffic enforcement.

By Keya Li, Jahnavi Malagavalli, Lamha Goel, Tong Wang, Kara M. Kockelman
arXiv AI
Sep 25

Less is More: Encoder-only Audio-Visual Segmentation

The paper "Less is More: Encoder-only Audio-Visual Segmentation" introduces EASE, an encoder-only model for Audio‑Visual Semantic Segmentation (AVSS). EASE achieves state‑of‑the‑art accuracy while running at up to 365 FPS—about three times faster than previous Transformer‑based AVSS models—and trains in under 11 GPU‑hours. The authors demonstrate that simpler, faster architectures can match or exceed the performance of more complex models across various backbones and resolutions.

By Ilpo Viertola, Vladimir Iashin, Sophie T\"otterstr\"om, Esa Rahtu
arXiv AI
Sep 25

GeoRefer-Bench: A Benchmark from Referring Pixels to Verifiable Geospatial Reasoning

GeoRefer-Bench is a new benchmark for verifiable geospatial referring segmentation that evaluates whether models correctly resolve spatial relations in overhead imagery. Each query is expressed as an executable logical form over a metric scene graph, and predictions are scored with Exact Query Success (EQS), requiring an exact match to the query’s referent set. The dataset contains 700 UAV scenes, 26,217 instances, 142,796 spatial relations, 20,916 executable queries across five reasoning levels, and additional paraphrases, unanswerable queries, counterfactual pairs, and leakage‑controlled splits.

By Shuaishuai Cao, Min Huang, Meng Tang, Xuan Liu, Youjin Wang, Hui Lin