Computer vision

Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.

3,375 stories · RSS feed

arXiv Computer Vision
Sep 28

AxonSynth: Domain-Randomized Synthetic Data for Zero-Shot 3D Axon Segmentation in Light-Sheet Microscopy

arXiv:2609.31431v1 Announce Type: new Abstract: Accurate segmentation of axons in 3D microscopy data is important for analyzing white-matter organization, but dense ground truth labels are expensive...

By Edward Gaibor, Kyriaki-Margarita Bintsi, Carmen Luz Leiva Ureta, Zayneb Bellatif, Chiara Maffei, Wenze Li, Elizabeth Hillman, Ya\"el Balbastre, Anastasia Yendiki
arXiv Computer Vision
Sep 25

GeoBlur: Epipolar Geometry Estimation from a Single Motion-Blurred Image

GeoBlur is a framework that estimates the fundamental matrix and relative camera pose from a single motion‑blurred image by exploiting blur artifacts as motion cues. It predicts visual correspondences between two time instances within the exposure window and solves the single‑frame epipolar geometry problem, yielding a fundamental matrix unique up to transposition due to time‑direction ambiguity. The method shows improved performance on synthetic and hybrid benchmarks and remains competitive on real motion‑blur data, also enabling downstream single‑frame motion segmentation.

By Bao-Long Tran, Cuong Le, Fredrik Viksten, Per-Erik Forss\'en
arXiv Computer Vision
Sep 25

Integrating Local Detail and Global Context: A Dual-Input Multi-Task Learning Framework for Bone Tumor Diagnosis

The paper introduces a dual‑input, multi‑task learning framework that jointly segments and classifies bone tumors by applying bidirectional cross‑modal attention between a lesion crop and the full radiograph. Using a YOLO‑based detector and a dual‑stream DenseNet121 architecture, the model fuses fine‑grained lesion detail with global anatomical context through a novel cross‑modal attention fusion strategy and hierarchical multi‑scale feature fusion. On the multi‑institutional Bone Tumor X‑ray Radiograph Dataset, the approach outperforms single‑input baselines, achieving a Dice coefficient of 0.896 and a macro‑averaged F1‑score of 0.928, with an AUC of 0.999 for malignant osteosarcoma.

By S. M. Nasif Uddin, Rusab Sarmun, Muhammad E. H. Chowdhury, Adam Mushtak, Israa Al-Hashimi, Sohaib Bassam Zoghoul
arXiv AI
Sep 25

Detecting Glaucoma Across Multi-ethnic Myopic and Non-Myopic Populations Using an Uncertainty-Aware Vision Transformer: A Multicentre Model Development and Validation Study

The study developed a Vision Transformer-based deep learning model with uncertainty estimation to detect glaucoma from colour fundus photographs across multi‑ethnic populations, including those with high myopia. Using 56,483 images for training, the model achieved an internal AUROC of 98.7% and maintained high performance (AUROC 86.4–99.6%) on 16 external datasets from eight countries. In high‑myopia eyes, the model outperformed ophthalmologists and matched specialists when full clinical data were available.

By Raghavan Lavanya, Yangqin Feng, Ten Cheer Quek, Quan V. Hoang, Linda Yi-Chieh Poon, Jost B. Jonas, Ya Xing Wang, Vinay Nangia, Jin Wook Jeoung, Sehie Park, SoYeon Kim, Benjamin Y Xu, Sreenidhi Iyengar Munimadugu, Paul Mitchell, Gerald Liew, Yanin Suwan, Jirayu Hong-amata, Sahil Thakur, Monisha E Nongipur, Tina Wong, Rahat Husain, Ng Si Rui, Yamon Syn, Phey Feng Lo, Nicholas Tan Yi Qiang, Shaista Hussain, Xiaofeng Lei, Zhi Da Soh, Marco Yu, Haslina Hamzah, Zizhou Wang, Yan Wang, Liangli Zhen, Xinxing Xu, Tien-Yin Wong, Tin Aung, Rachel S Chong, Yong Liu, Ching-Yu Cheng
arXiv Computation and Language
Sep 25

DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units

DiscoPhon is a multilingual benchmark designed to evaluate unsupervised phoneme discovery from discrete speech units. It includes 6 development and 6 test languages that cover a wide range of phonemic contrasts, and requires systems to generate discrete units mapped to a predefined phoneme inventory using only 10 hours of speech from an unseen language. The benchmark assesses unit quality, recognition, and segmentation, and provides four pretrained multilingual HuBERT and SpidR baselines that demonstrate current models can produce units that correlate well with phonemes, though performance varies across languages.

By Maxime Poli, Manel Khentout, Angelo Ortiz Tandazo, Ewan Dunbar, Emmanuel Chemla, Emmanuel Dupoux
arXiv Computer Vision
Sep 25

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

RGBD20K is a new large-scale RGB‑D semantic segmentation dataset featuring 20,000 image pairs and 160 fine‑grained categories, surpassing existing benchmarks like NYUv2 and SUN RGB‑D in both scale and semantic diversity. The dataset provides high‑fidelity annotations obtained through rigorous re‑evaluation and correction of prior labels, ensuring a clean ground‑truth foundation. Additionally, the authors introduce a score‑purified fusion (SPF) method that achieves state‑of‑the‑art performance across evaluated benchmarks, demonstrating the value of high‑quality multimodal information.

By Shaohua Dong, Zexuan Meng, Haiyan Sun, Bing Fan, Cuicui Zhang, Dylan Joseph, Kewei Sha, Yunhe Feng, Heng Fan
arXiv Computer Vision
Sep 25

Context-aware Skin Cancer Epithelial Cell Classification with Scalable Graph Transformers

The paper introduces scalable Graph Transformers for classifying healthy versus tumor epithelial cells in whole-slide images of cutaneous squamous cell carcinoma. By constructing a full‑WSI cell graph and incorporating morphological, texture, and neighboring cell class features, the proposed SGFormer and DIFFormer models outperform traditional image‑based methods, achieving balanced accuracies above 85% on single‑WSI tests and 83.6% on multi‑WSI evaluations. The study demonstrates that preserving tissue‑level context through graph representations improves classification of morphologically similar cell types.

By Lucas Sanc\'er\'e, No\'emie Moreau, Katarzyna Bozek
arXiv Computer Vision
Sep 25

FoCal: Frequency-Oriented Cross-Modal Interaction and Spectral Calibration for Aerial Visible-Infrared Object Detection

FoCal is a new frequency‑oriented framework for aerial RGB–IR object detection that explicitly models cross‑modal interaction across different frequency components. It introduces a Frequency‑Aware Dual‑Domain Calibration module to consolidate low‑frequency structural cues while preserving high‑frequency modality‑specific details, and a Discrepancy‑Guided Spectral Modulation module that adaptively enhances, preserves, or attenuates the joint spectrum based on confidence‑weighted amplitude discrepancies. Experiments on DroneVehicle, ESCVehicle, and ATR‑UMOD show FoCal achieving high mAP scores (83.5%, 54.8%, 64.6%) with only 3.0 M parameters and 113.6 FPS, demonstrating a strong accuracy–efficiency trade‑off.

By Ben Liang, Chao Sui, Junqi Bai, Yuan Liu, Chunlai Li, Xiubao Sui, Qian Chen
arXiv Computer Vision
Sep 25

When Misalignment Becomes Supervision: Structured Label Noise in Supervised Synthetic CT Generation

The paper examines how residual misalignments from registration procedures introduce structured label noise in supervised synthetic CT (sCT) generation. It shows that voxel‑wise metrics are heavily influenced by the consistency between training and evaluation registrations, and that training with anatomically consistent registrations reduces variability and improves robustness. Introducing a perceptual loss based on a pretrained Segment Anything encoder yields sharper, more anatomically coherent sCT and highlights the need for anatomy‑oriented evaluation.

By Valentin Boussot, Cedric Hemon, Caroline Lafond, Jean-Claude Nunes, Jean-Louis Dillenseger
arXiv Computer Vision
Sep 25

Match4Annotate: Cross-Video Annotation Transfer in Ultrasound via Implicit Feature Flow-Guided Matching

Match4Annotate is a test‑time framework that transfers user‑specified annotations from a labeled ultrasound video to an unlabeled target video without requiring manual initialization. It uses a spatiotemporal implicit feature representation, a continuous implicit feature flow for alignment, and flow‑guided annotation transfer to unify sparse point and dense mask transfer. The method achieves state‑of‑the‑art performance on four clinical ultrasound datasets, outperforming dense feature‑matching baselines and one‑shot segmentation methods, and works without task‑specific training on a single consumer GPU.

By Zhuorui Zhang, Roger Pallar\`es-L\'opez, Praneeth Namburi, Brian W. Anthony
arXiv AI
Sep 25

SARFusion: Scene-Aware Routing Fusion for Robust Camera-LiDAR 3D Object Detection

SARFusion introduces a scene-aware routing approach for camera‑LiDAR 3D object detection, decoupling object‑query decoding into separate camera, LiDAR, and fusion branches. By estimating a global scene reliability prior and incorporating object‑level evidence, each query is routed to the most suitable branch, reducing cross‑modal interference. The method achieves strong performance on the nuScenes test set (72.5 mAP, 74.4 NDS) and demonstrates robustness to sensor corruptions and environmental changes.

By Yuting Zhao, Ziyi Zheng, Shuxiao Li
arXiv AI
Sep 25

TOLA: Text-aware One-Step Latent Adaptation for Diffusion-based Text Image Super-Resolution

TOLA is a diffusion‑based text image super‑resolution method that eliminates iterative image‑text diffusion by using a one‑step latent adaptation framework. It employs a confidence‑weighted text conditioning module to build a reliable semantic condition and a lightweight latent residual correction module to fix structured residual errors, thereby preserving text fidelity. Experiments show TOLA outperforms existing diffusion‑based TSR methods, achieving at least 2.72 dB higher PSNR on the CTR‑TSR‑Test benchmark.

By Yike Xu, Yue Shi, Yong Guo, Jiezhang Cao
arXiv Computer Vision
Sep 25

BiCC: Bidirectional Connected-Component Loss for Instance-Aware Segmentation

The paper introduces BiCC, a bidirectional connected-component loss that pairs annotation- and prediction-derived partitions to score predicted components on their own scale. By deriving instances from predictions, BiCC directly penalizes false-positive components regardless of size, allowing a balance parameter to control the lesion-wise precision–recall trade-off. Across five datasets, BiCC outperforms existing instance-aware losses such as CC-DiceCE and blob loss in lesion-wise F1, and improves over DiceCE on multiple datasets.

By Luc Bouteille, Frederic Jonske, Jens Kleesiek, Alexander Jaus
arXiv Computer Vision
Sep 25

Does DCGAN-Based Synthetic Augmentation Improve Brain Tumor MRI Classification? An Empirical Study

This study examined whether augmenting brain tumor MRI datasets with class‑specific DCGAN‑generated images improves classification performance. Using 7,200 scans across four tumor categories, a Swin Transformer classifier trained on real images alone achieved 96% accuracy, identical to the model trained with 500 synthetic images per class. Metrics such as macro F1 and ROC‑AUC showed no improvement, and FID scores indicated substantial distributional differences between real and synthetic images.

By Irhum Jawad Khan, Talha bin Aslam
arXiv Machine Learning
Sep 25

Towards Deployable Underwater Vessel Classification

The paper presents a compact underwater acoustic classification framework that integrates multi-representation feature engineering, temporal statistical pooling, and lightweight convolutional architectures for acoustic time-frequency and cochlear representations. Experiments on the ShipsEar dataset show a two-layer CNN achieving a macro F1 of 0.9918 and an RBF-SVM reaching 0.9883, but recording provenance issues limit verification of generalisation. When evaluated on the DeepShip dataset with recording-level partitioning, a 157K-parameter CNN attains a macro F1 of 0.7226, while a larger ResNet18 does not improve validation performance, underscoring the need for representation-aware design and rigorous evaluation for deployable systems.

By Abishek Soti, Thura Pyae Sone, Naqib Ibnul, Htoo Htet Aung, Henry Zhong, Gregory Cohen, Ying Xu
arXiv Computer Vision
Sep 25

Lightweight Vision Transformer-Based U-Net for Brain Tumor Segmentation from MRI

The paper introduces a lightweight Vision Transformer‑based U‑Net for brain tumor segmentation from MRI, combining U‑Net’s hierarchical feature extraction with a compact ViT bottleneck to capture both local and global context. With only 2.6 million trainable parameters, the model achieves a mean Intersection over Union of 0.8100 and a Dice score of 0.8446 on the TCGA LGG dataset, surpassing the baseline U‑Net by 3.75% and 3.15% respectively. Extensive quantitative and qualitative analyses, including confusion matrices, precision‑recall curves, and tumor size dependency studies, demonstrate the method’s effectiveness and robustness.

By Sheekar Banerjee, Md. Srabon Chowdhury, Md. Mahbub Hasan Akash, Ishtiak Al Mamoon
arXiv AI
Sep 25

UNWIND: Any-Length Facial Video for Stress Detection without Temporal Windowing

UNWIND is a facial‑video framework that detects stress by treating an entire recording as a single input, avoiding the need for temporal windowing or segmentation. It folds the video’s temporal dimension into the channel dimension of a 2‑D spatial representation and processes it with an asymmetric‑attention architecture. Experiments on a 58‑subject stress dataset show that using all 3,600 frames (stride τ = 1) yields a 69.73 % accuracy, comparable to the best 70.02 % accuracy at τ = 15, while computational cost varies from 12.48 to 348.78 GFLOPs.

By Stefanos Gkikas, Christian Arzate Cruz, Eric Nichols, Giorgos Giannakakis, Randy Gomez
arXiv Computer Vision
Sep 25

Retrieve-to-Localize: Bridging Large Language Models and LiDAR Geometry for Spatial Grounding

The paper introduces a method that combines large language models (LLMs) with LiDAR geometry to answer complex spatial questions by grounding targets directly in LiDAR point clouds. It presents the SpatialLiDAR-QA dataset for relational grounding tasks and the SpatialLiDAR-LM model, which aligns LiDAR features with an LLM to retrieve and refine target coordinates. Experiments show significant gains over existing LiDAR–language models and multi‑camera vision‑language models in precise coordinate prediction.

By Byounggun Park, Giyong Moon, Jusung Kim, Soonmin Hwang
arXiv AI
Sep 25

Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams

The study evaluates large language model (LLM) graders on two computer‑science exams, testing 171 configurations of closed‑ and open‑weights models. While the best LLM configuration achieved a mean absolute error of 1.64/35—better than the 2.61/35 error between two human graders—its performance was highly sensitive to the prompt. A short "strict grader" preamble caused most open‑weight models to exceed acceptable error thresholds or stop grading entirely, whereas fine‑tuning with a single LoRA adapter restored parity with human graders and reduced sensitivity to harsh prompts.

By Ali Habibullah, Yazan Alshoibi, Mohammad Alshiekh, Salman Khan, Naeemullah Khan
arXiv AI
Sep 25

When Explanations Cannot Be Read: Measuring and Correcting SHAP and LIME Rendering for Right-to-Left Languages

The paper "When Explanations Cannot Be Read: Measuring and Correcting SHAP and LIME Rendering for Right-to-Left Languages" identifies that standard SHAP and LIME visualizations, designed for left‑to‑right scripts, fail to display attribution values correctly for right‑to‑left languages such as Urdu, Arabic, Persian, and Hebrew. It introduces SHAP‑RTL, a rendering layer that preserves the original attribution values while correcting reading direction, script shaping, and font selection for each language. The authors evaluate SHAP‑RTL on hate‑and‑offensive‑language datasets using TF‑IDF and logistic regression, showing that default rendering yields high character error rates, while SHAP‑RTL maintains correct visualizations across all tested languages.

By Rameesha Zia, Muhammad Shahid Iqbal Malik