Computer vision

Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.

3,375 stories · RSS feed

arXiv AI
Sep 25

QINA: Quantum-Inspired Nonlinear Adapters for Pretrained Vision Models

The paper introduces Quantum-Inspired Nonlinear Adapters (QINA), compact modules that apply learnable trigonometric feature lifting followed by bounded nonlinear aggregation to pretrained vision models. QINA enables structured oscillatory basis functions with a norm-dependent Lipschitz bound, allowing spectral reshaping of representations without expanding the receptive field or significantly increasing parameters. Experiments on natural and medical imaging tasks show that QINA consistently outperforms identity baselines, fixed Fourier mappings, and parameter-matched generic adapters, demonstrating that geometry- and spectrum-aware adaptation is crucial for effective frozen-backbone transfer learning.

By Mostafa Mehdipour Ghazi
arXiv AI
Sep 25

Mind What Matters for Reasoning: Aligning Cross-Modal Attention via Selective Probability Mass Concentration

The paper introduces Selective Probability Mass Concentration (sPMC), a training framework that strengthens implicit visual grounding in multimodal large language models by selectively regularizing attention heads most responsive to visual evidence. sPMC treats attention over visual tokens as a spatial probability distribution and encourages mass to concentrate on semantically relevant regions using segmentation-derived priors, while leaving other heads unconstrained. Across six multimodal benchmarks, sPMC yields an average zero‑shot improvement of 3% and gains up to 11.3% for various models by regularizing only 3%–15% of their attention heads.

By Jiaqi Deng, Zonghan Wu, Zhan Heng, Xiaoshui Huang, Huan Huo, Guandong Xu
arXiv Machine Learning
Sep 25

Albireo: Adaptive, Energy-Efficient Inference Framework for Video Object Detection on the Edge

Albireo is an adaptive, energy‑efficient inference framework for video object detection on edge devices that wraps existing detectors without modification. It uses a 10‑dimensional Kalman filter per active object to decide when to skip detector calls, predicting bounding boxes on skipped frames at near‑zero GPU cost. Evaluated on BDD100K with YOLO and RF‑DETR detectors on NVIDIA Jetson AGX Thor and Orin, Albireo maintains AP@50 within ±1.2 pp of full‑frame inference while reducing energy consumption by 12.1–17.6 % and improving accuracy for some models.

By Amir Taherin, Jos\'e Cano, Bin Ren, Yanzhi Wang, David Kaeli
arXiv Machine Learning
Sep 25

Named Entity Recognition using Sliding Window Approach

The paper presents an inference-only pipeline that extends the frozen NER model MahaNER‑BERT to document‑level prediction using overlapping sliding windows, eliminating the need for retraining or architectural changes. The approach is evaluated on six document‑level corpora derived from the MahaNER test set, employing two repetition strategies (Normal Repeat and Random Repeat) at three length levels and various window configurations. Results show the model maintains a macro F1‑score of up to 0.8902 with minimal variation, outperforming non‑windowed methods by avoiding boundary‑fragmentation errors and achieving more stable document‑level performance.

By Hariom Ingle, Ronit Ghode, Ishwari Gondkar, Jidnyasa Harad, Ravindra Murumkar, Raviraj Joshi
arXiv Machine Learning
Sep 25

Calpric: Inclusive and Fine-grain Labeling of Privacy Policies with Crowdsourcing and Active Learning

Calpric is a system that combines automatic text selection, segmentation, active learning, and crowdsourced annotation to create a large, balanced training set for privacy policy classification. By simplifying the labeling task, it enables untrained crowd workers to match the performance of trained annotators and reduces inter‑annotator disagreement, cutting labeling costs. The approach yields a dataset of 16,000 policy text segments across nine data categories and produces models that deliver accurate, fine‑grained labels at a cost of roughly $0.92–$1.71 per segment.

By Wenjun Qiu, David Lie, Lisa Austin
arXiv Computation and Language
Sep 25

PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs

PTC-Bias is a two-stage framework that improves contextual biasing in speech large language models by using phoneme-level temporal competition. In the first stage, PTC Retrieval performs frame-synchronous phoneme decoding to generate a compact shortlist of bias words and their speech intervals. The second stage, PTC Correction, applies a local competition between retrieved candidates and mismatched transcript spans within those intervals, reducing near-homophone and word-segmentation errors without extra SpeechLLM passes. Experiments on LibriSpeech demonstrate consistent gains across two SpeechLLMs, with PTC-Bias reducing B-WER by up to 23.9% relative to CTC-Filter while keeping U-WER nearly unchanged.

By Zhiqi Ai, Han Cheng, Shiyi Mu, Yongjin Zhou, Shugong Xu
arXiv Computer Vision
Sep 25

Bridging the Inter-Domain Gap through Low-Level Features for Cross-Modal Medical Image Segmentation

The paper introduces LowBridge, a method for cross‑modal medical image segmentation that leverages shared low‑level features such as edges between MRI and CT scans. It trains a generative model to reconstruct source‑modality images from edge maps and then trains a segmentation network on these generated images. At test time, edge features from target‑modality images are fed into the generative model to produce source‑style images, which are segmented by the pretrained network, achieving state‑of‑the‑art results on multiple public datasets.

By Pengfei Lyu, Pak-Hei Yeung, Jing Xia, De Hu, Xiaosheng Yu, Jianning Chi, Chengdong Wu, Jagath C. Rajapakse
arXiv AI
Sep 25

A Multimodal 3D Foundation Model for Light Sheet Fluorescence Microscopy Enables Few-Shot Segmentation, Classification, and Deblurring

The paper presents a 3D foundation model for light sheet fluorescence microscopy (LSM) that is pretrained on a large curated set of 3D images from various organisms, stains, and imaging protocols. By jointly optimizing for masked reconstruction and image‑text alignment, the model learns transferable volumetric representations that dramatically reduce the need for annotated data. The pretrained backbone enables efficient few‑shot adaptation to downstream tasks such as segmentation, classification, and deblurring, consistently outperforming baselines according to standard metrics and expert evaluation.

By Adina Scheinfeld, Haotan Zhang, Shang Mu, Rudolf L. M. van Herten, Lucas Stoffl, Ali Erturk, Zhuhao Wu, Johannes C. Paetzold
arXiv Computation and Language
Sep 25

Small yet Assistive: Spatially-Aware Post-Training for Low Vision

The paper introduces Smol‑VL‑BLV, a compact vision‑language model designed for blind and low‑vision users. It employs a 500M decoder transformer with teacher‑student distillation and Group Relative Policy Optimization to add spatial detail, directional cues, and hazard detection to post‑training. After a lightweight finetuning step, the model achieves significant gains on spatial, social, OCR, and VQA benchmarks while remaining under 450 MB and running entirely offline on a mid‑range Android phone.

By Rishabh Choudhary, Shreyansh Raj, Umesh Goyal, Shubh Kashyap, Shrestha Kumar, Sushovan Jena, Komal Kumar, Hisham Cholakkal, Aditya Nigam
arXiv Computer Vision
Sep 25

Comparing YOLOv11 and YOLOv8 for instance segmentation of occluded and non-occluded immature green fruits in complex orchard environment

This study evaluates the instance‑segmentation performance of YOLOv11 and YOLOv8 on immature green apples in orchard settings. YOLO11n‑seg achieved the highest mask precision (0.831), while YOLO11m‑seg and YOLO11l‑seg excelled in non‑occluded and occluded fruitlet segmentation. YOLOv8n, however, outperformed the YOLO11 series in inference speed, reaching 3.3 ms compared to 4.8 ms for the fastest YOLO11 model.

By Ranjan Sapkota, Manoj Karkee
arXiv Computer Vision
Sep 25

SplatLabel: Pseudo-Labelling through 4D Gaussian Splatting

SplatLabel is an automated pipeline that uses a 4D Gaussian representation to generate LiDAR segmentation and semantic occupancy grids with predictive confidence. It models dynamic scenes through an explicit temporal manifold, tracking moving actors without requiring pre‑annotated 3D bounding boxes. By integrating 360‑degree LiDAR depth maps and distilling soft probabilities from 2D models, it resolves semantic ambiguities over time and space, and evaluates pseudo‑labels via a selective classification framework that balances precision and recall.

By Nitya Nanvani, Andras Palffy, Holger Caesar
arXiv Computer Vision
Sep 25

Training-Free Hold-Usage Detection in Sport Climbing with Foundation Pose Models

The paper presents a training‑free method for detecting which holds a climber uses in sport climbing videos by leveraging a frozen foundation pose model (Sapiens) that provides fingertip and toe keypoints. Using a simple proximity test, mutual exclusion, and a temporal‑persistence rule, the approach achieves high F_1 scores (up to 90.2%) on the Way Up dataset without any climbing‑specific training, outperforming repurposed pose pipelines. The resulting automatic predictions enable accurate coaching statistics, such as climb time and pace, with Pearson correlations of 1.00 and 0.94 respectively.

By Abu Bakar, Abdullah Aftab, Amir Hamza
arXiv AI
Sep 25

CoSWA-YOLOv12: Scale-Invariant Tiny Object Detection and Segmentation of Malaria Parasites

CoSWA-YOLOv12 is a compact YOLOv12 instance‑segmentation detector designed to improve detection of tiny malaria parasites in microscopy images. It introduces a Cooperative Scale‑adaptive Wasserstein Assignment that applies a Wasserstein distance‑based label assignment inversely proportional to object size, a wavelet detail residual, and a min‑max Gaussian regression loss, all of which preserve pretrained weights. On a five‑class Rwandan thick‑smear dataset, the method raises P. falciparum recall from 0.63 to 0.74, increases mAP@50 from 0.73 to 0.81, and reduces missed detections from 38% to 15%.

By Ahmed Tahiru Issah, Carine Mukamakuza
arXiv Machine Learning
Sep 25

UltraBench 2: Towards Robust Evaluation of Vision Foundation Models on Ultrasound

UltraBench 2 is a new benchmark designed to evaluate vision foundation models on ultrasound images, addressing the lack of standardized tests in this area. It covers a wide range of anatomical structures and tasks, emphasizing reproducibility and ease of use. The authors compare existing models, finding that ultrasound-specific pretraining still outperforms on classification, while general-purpose models have matched performance on segmentation.

By Ashwath Radhachandran, Adam Tupper, Christian Gagn\'e, William Speier
arXiv Computer Vision
Sep 25

OncoVision: Integrating Mammography and Clinical Data through Attention-Driven Multimodal AI for Enhanced Breast Cancer Diagnosis

OncoVision is a privileged‑information training framework that learns from mammography images and clinical data during training but performs inference using only mammographic images. It employs an attention‑based encoder‑decoder to jointly segment masses, calcifications, axillary findings, and breast tissue, and predicts ten structured clinical features such as BI‑RADS. Two late‑fusion strategies (Independent and Dependent) integrate imaging, radiomic, and clinical information to improve diagnostic precision, and a retrospective multi‑reader study showed higher diagnostic confidence, reduced reading time, and segmentation accuracy comparable to or better than radiologists.

By Istiak Ahmed, Galib Ahmed, K. Shahriar Sanjid, Md. Tanzim Hossain, Md. Nishan Khan, Md. Misbah Khan, Md. Arifur Rahman, Sheikh Anisul Haque, Sharmin Akhtar Rupa, Mohammed Mejbahuddin Mia, Mahmud Hasan Mostofa Kamal, Md. Mostafa Kamal Sarker, M. Monir Uddin
Hugging Face Trending Papers
Sep 24

Less is More: Encoder-only Audio-Visual Segmentation

The paper introduces EASE, an encoder‑only model for Audio‑Visual Semantic Segmentation that eliminates redundant components found in prior Transformer‑based approaches. EASE achieves state‑of‑the‑art accuracy while running at up to 365 FPS—three times faster than previous models—and can be trained in under 11 GPU‑hours. The authors provide code, weights, and samples, positioning EASE as a scalable foundation for future research and real‑time applications.

Hugging Face Trending Papers
Sep 24

RGBD20K: A Large-Scale Benchmark for RGB-D Semantic Segmentation

RGBD20K is a new large-scale RGB‑D dataset designed to advance semantic segmentation research. It contains 20,000 image pairs annotated with 160 fine‑grained categories, far exceeding the diversity of existing benchmarks such as NYUv2 and SUN RGB‑D. The authors also provide high‑fidelity annotations and introduce a score‑purified fusion (SPF) method that achieves state‑of‑the‑art results on multiple benchmarks.

arXiv Machine Learning
Sep 24

M3D-Net: Hierarchical Coordination of Spatial Context, Feature Reuse, and Differential Attention for Mammography Classification

M3D‑Net is a mammography encoder that hierarchically coordinates multi‑scale coordinate attention, bounded dynamic feature reuse, and differential attention through resolution‑aware operator placement. It preserves earlier features within stages, integrates local and global context via coordinate‑aware aggregation, and applies differential attention at coarse resolutions. In image‑only classification on AISSLab mammography and an adapted image‑clinical model on BrEaST ultrasound, M3D‑Net achieves the highest validation accuracy and lowest endpoint cross‑entropy loss compared to EdgeNeXt, RepViT, and TransXNet, with accuracies of 97.78% and 80.39% respectively.

By Zheng Yu, Xinhang Li, Jiabao Gao, Boyang Wang, Xiang Li
arXiv Computer Vision
Sep 24

nnFoundation: 3D Foundation Models for Radiology

nnFoundation introduces complementary convolutional and transformer-based 3D foundation models for radiology, trained on 2.1 million CT, MRI, and PET volumes from 125 datasets. The models are evaluated on 108 tasks—including segmentation, detection, classification, report generation, and image retrieval—under domain shift, low-data, and low-compute scenarios, consistently outperforming prior 3D foundation models and training from scratch. Performance varies by task type, with convolutional models excelling at spatially localized tasks and transformer models at global semantic reasoning, and dynamic alignment with dataset characteristics further enhances transferability.

By Constantin Ulrich Harsy, Tassilo Wald, Karol Gotkowski, Yannick Kirchhoff, Marcel Knopp, Maximilian Rokuss, Elisa Stegmeier, Philipp Schader, Dasha Trofimova, Raphael Stock, Kim-Celine Kahl, Stephen Schaumann, Selen Erkan, David Zimmerer, Stefan Denner, Moritz Langenberg, Sebastian Ziegler, Katharina Eckstein, Maximilian Fischer, Jonathan Suprijadi, B\'alint Kov\'acs, Benjamin Hamm, Anand Deshpande, Dimitrios Bounias, Nico Disch, Shuhan Xiao, Jessica K\"achele, Jan Sellner, Rajesh Baidya, Jeremias Traub, Lars Kr\"amer, Maximilian Zenk, Tim R\"adsch, Stefan Dvoretskii, Robin Peretzke, Jonathan Deissler, Alexandra Ertl, Partha Ghosh, Kris Dreher, Stefan Dinkelacker, Annika Reinke, Evangelia Christodoulou, Numan Saeed, Yoland Savriama, Santiago Estrada, David K\"ugler, Laura Alexandra Daza Barragan, Cristina Isabel Gonzalez Osorio, Jan Peeken, Michael Baumgartner, Marvin Teichmann, Guillaume Chabin, Matthias Kirchler, Valentin Koch, for the ALFA study, Markus Hohenhaus, Dimitri Koslov, Nina Decker, Mohammad Yaqub, Arnd Heuser, Martin Reuter, Julia A. Schnabel, Tobias Heimann, Florin Ghesu, Paul Brachmann, Claus P. Heu{\ss}el, Alexander Radbruch, Gianluca Brugnara, Aditya Rastogi, Martha Foltyn-Dumitru, Heinz-Peter Schlemmer, Ignaz Reicht, Julius C. Holzschuh, Michael Bach, Bram Stieltjes, Kai Schlamp, Lena Maier-Hein, Marco Nolden, Ralf Floca, Paul F. J\"ager, Philipp Vollmuth, Fabian Isensee, Klaus H. Maier-Hein