Computer vision

Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.

3,375 stories · RSS feed

arXiv AI
Sep 28

The Right Information Extraction Pipeline Depends on the Document: Accuracy-Energy Trade-offs for Small, Local Models

The study investigates how the choice between processing page images or parsed text in an information extraction pipeline depends on the document’s layout, focusing on privacy‑sensitive, on‑premise scenarios with small models (≤8 B parameters). It evaluates accuracy and energy consumption across input representations, model families, and inference settings, finding that batching dramatically reduces energy use, FP8 quantization offers modest savings, and neural OCR is far more energy‑intensive than classical OCR. The optimal representation varies: vision‑language models excel on layout‑rich documents, while small text‑only models with a cheap parser perform best on near‑plain‑text contracts, achieving higher accuracy and lower energy than any vision‑language setup.

By Christoph Walser, Mauricio Fadel Argerich, Jonathan F\"urst
arXiv AI
Sep 28

MedTokenBudget: Lesion-Preserving Token Routing for Dermoscopic Image Classification

MedTokenBudget introduces a supervised token routing framework for Vision Transformers applied to dermoscopic image classification. Its Lesion-Aware Token Scoring (LATS) module combines attention entropy, feature norm, and local feature contrast to select the top‑K patches under a target budget, trained with curriculum learning, diversity regularization, attention distillation, and lesion‑mask supervision. On the ISIC 2019 dataset, mask‑supervised LATS outperforms Random and ToMe at headline budgets while retaining more lesion patches.

By Zhexiang Li
arXiv AI
Sep 28

Combining General and Domain-Specific Pretext Tasks for Brain MR Image Segmentation

The paper investigates combining a domain‑specific self‑supervised task—voxel‑level brain age prediction—with a general task—image inpainting—to pretrain models for brain MRI segmentation. A multitask pretraining framework jointly optimizes both objectives, yielding representations that outperform single‑task pretraining and training from scratch on three segmentation benchmarks (multiple sclerosis lesions, ischemic stroke lesions, and cortical structures). The study demonstrates that integrating domain‑specific and general self‑supervised tasks benefits the development of generalizable neuroimaging foundation models.

By Tasneem Nasser, Susanne Schmid, Roberto Souza, Naser El-Sheimy
arXiv AI
Sep 28

Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models

The paper introduces LIRSeg, a method that replaces explicit Chain-of-Thought reasoning in multimodal large language models with a compact set of learnable latent tokens for reasoning segmentation. LIRSeg is trained in two stages—spatial alignment and GRPO—while employing extreme-advantage sampling, decoupled exploration-stability updates, and latent diversity amplification to enhance token informativeness. Experiments show that LIRSeg improves segmentation accuracy and reasoning efficiency, achieving significant gIoU gains over the VisionReasoner baseline and reducing reasoning tokens by about 16×.

By Tianhang Guo, Yulin He, Wei Chen, Wenjuan Zhou, Yuhang Li, Xinbiao Gan
arXiv AI
Sep 28

Adaptive Pilot Selection for Unified Semantic Communication and Semantic Sensing in ISAC

The paper introduces SemISAC, a unified framework that integrates semantic communication and semantic sensing into a single dual‑function waveform. It employs a joint semantic encoder to extract task‑specific information for both communication and sensing, and uses an adaptive pilot configuration to balance channel estimation and sensing needs. In vehicular scenarios, SemISAC achieves segmentation accuracy comparable to dedicated semantic communication systems while outperforming conventional and semantic baselines in target recognition and range estimation.

By Muhammad Abubakar Rashid, Muhammad Hannan Akram, Haejoon Jung, Syed Ali Hassan
arXiv AI
Sep 28

UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound

UltraG-Bench is a large‑scale, multi‑task benchmark designed to evaluate pixel‑level evidence grounding in ultrasound images. It comprises 40 public segmentation datasets covering 13 anatomical categories and includes three progressive tasks—instruction‑guided segmentation, evidence‑grounded VQA, and evidence‑grounded report generation—with a total of 736,726 annotations. Evaluation of 14 state‑of‑the‑art models shows a significant gap between semantic understanding and fine‑grained pixel‑level localization, and the authors propose UltraG‑Agent, which combines a vision‑language model with the ultrasound‑specific segmentation model UltraSAM to improve both semantic prediction and visual grounding.

By Quanhao Zhu, Bo Xu, Rui Lin, Chenyuan Wang, Yu Shao, Boling Zhu, Jiuyan Sun, Liang Zhao, Hongfei Lin, Feng Xia
arXiv AI
Sep 28

DepthEvidence: Unifying Metric Depth Prediction and Geometric Reasoning in Multimodal Language Models

DepthEvidence is a 4B multimodal language model that integrates dense metric depth predictions into language generation. It employs a camera‑conditioned decoder to produce full‑resolution depth maps and a dense‑to‑language interface that converts these predictions into object‑aligned geometry tokens. The model is trained with geometric supervision and instruction tuning, and it sets new state‑of‑the‑art results on a Depth‑VQA benchmark and on instance‑level metric depth estimation across nine datasets.

By Jiangning Wei, Yuan Yao, Miaomiao Cui, Mingsheng Li, Humen Zhong, Shuai Bai, Zhibo Yang
arXiv AI
Sep 28

Teacher-Anchored Selection of Post-Training Quantized Models under Domain Shift

The paper investigates how to choose the best quantized model from a family of compressed versions when target labels are scarce or unavailable. It finds that a simple rule based on minimum teacher distortion consistently selects the same eight‑bit, per‑channel, unclipped configuration, though this does not minimize empirical target cross‑entropy. The study also shows that confidence‑based estimators perform poorly in overconfident regimes, while output‑distribution estimators can outperform the teacher in some architectures, and that combining distortion with a supervised term can improve selection. Across 134 candidate families, teacher‑anchored selection reduces mean regret with very few labels, though the benefit diminishes after about 25 labels.

By Alejandro Rodriguez Dominguez, Muhammad Shahzad, Xia Hong
arXiv AI
Sep 28

ReG-SAM: Reference Graph-Driven SAM for 2D Foundational Vessel Segmentation

ReG-SAM is a SAM-based framework designed for 2D vessel segmentation in medical images. It introduces reference graph prompt embeddings (GPEs) and vascular prototype embeddings (VPEs) to capture global spatial and fine-grained modality-specific vessel features, respectively. By building a modality-wise vascular database and learning these embeddings from reference masks, ReG-SAM consistently outperforms existing baselines across 19 datasets, especially on thin vessels.

By Donghang Lyu, Zichen Zhang, Oleh Dzyubachyk, Marius Staring
arXiv AI
Sep 28

Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers

The paper "Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers" reports that vision‑language models (VLMs) struggle to read images containing two overlapping text layers—one with sharp contour lines and one with soft shading. Using the DecoyBench dataset of 300 such images, the authors evaluated six closed‑source VLMs under naive and guided prompting at high and low resolutions. While humans could read both layers accurately, the models reliably read only the contour layer at high resolution and failed to extract the shading layer; at low resolution, neither the models nor humans could read the contour layer, but the shading layer remained readable. "whyItMatters":"The study highlights a consistent limitation of current VLMs in handling typographic structures with multiple spatial frequency layers, underscoring their vulnerability to typographic attacks and the need for more robust text‑recognition capabilities."

By Mert \.Incidelen, Yamen Kashkash, Asya Berker, Murat Aydo\u{g}an
arXiv Computer Vision
Sep 28

Double-stream registration with pyramid fusion for HDR video with alternating exposures

The paper introduces a new HDR video reconstruction framework that uses a double-stream registration approach combined with pyramid fusion. It processes three consecutive frames by computing optical flow relative to the central frame and applying a midpoint displacement strategy to address severe overexposure. The resulting radiance and low‑dynamic‑range images are merged in a pyramid fusion stage to produce the final HDR output, and experiments show the method outperforms existing state‑of‑the‑art techniques.

By Onofre Martorell, Ivan Pereira-S\'anchez, Antoni Fuentes, Antoni Buades
arXiv Computer Vision
Sep 28

A Multi-Stage Framework for Kuzushiji Character Recognition in Japanese Historical Documents

The paper presents a multi‑stage framework for recognizing Kuzushiji characters in Japanese historical documents. It combines character detection, cropping, classification, reading‑order reconstruction via adaptive column clustering, and large‑language‑model‑based post‑OCR correction. The authors also augment data synthetically, correct dataset annotations, and introduce new test sets, achieving significant character error rate reductions on real, synthetic, and out‑of‑domain data.

By Rui-Yang Ju, Kohei Yamashita, Hirotaka Kameko, Shinsuke Mori
arXiv Computer Vision
Sep 28

ManiVid: Unified and Explainable Forensic Analysis of Manipulated Videos

ManiVid introduces a unified forensic analysis framework for manipulated videos, combining forgery detection, artifact grounding, and anomaly explanation. The authors release ManiVid-38K, a large dataset of 19K real‑fake video pairs with authenticity labels, forgery masks, and explanations, and a benchmark ManiVidBench with 1K balanced pairs. ManiVidLens, the proposed model, outperforms existing methods in artifact grounding and anomaly explanation while matching state‑of‑the‑art detection accuracy.

By Hengrui Kang, Zhonghao Yan, Yuxuan Yang, Ruoyan Jing, Yuncheng Guo, Hao Chen, Kongming Liang, Zhanyu Ma, Conghui He, Weijia Li
arXiv Computer Vision
Sep 28

PICO: Projection-Informed Consistency Optimisation for 6DoF Surgical Tool Pose Estimation

The paper introduces PICO, an end-to-end trainable model for 6DoF surgical tool pose estimation that uses multi-task learning to predict segmentation, depth, and pose parameters. It incorporates two geometry-aware proxy tasks—a projection loss and a point-to-point loss—to enforce consistency in 2D and 3D spaces, improving accuracy and robustness. Evaluated on the SurgRIPE dataset, PICO achieves strong performance, ranking second in rotation accuracy and maintaining competitive translation results, especially under occlusion.

By Lucy Fothergill, Pietro Valdastri, Dominic Jones, Duygu Sarikaya
arXiv Computer Vision
Sep 28

TRACKGRAPH: Online Open-Vocabulary 3D Scene Graphs via Image-Space Tracking

TRACKGRAPH is an online open‑vocabulary 3D mapping system that tracks 2D masks in the image stream before fusing them into a class‑agnostic 3D segment layer within a hierarchical scene graph. It uses FastSAM and CLIP for sparse keyframes, DINOv3 for dense mask propagation, and compact multi‑view CLIP embeddings for open‑vocabulary retrieval. The method outperforms state‑of‑the‑art mapping techniques on Replica, ScanNet++, and HM3D, achieving higher synonym frequency, faster processing, and lower GPU memory usage, and has been deployed on quadruped robots and drones at real‑time rates.

By Peder Borge Hellesylt, Albert Gassol Puigjaner, Kostas Alexis, Annette Stahl
arXiv Computer Vision
Sep 28

Band-Selection Stability and Semantic Segmentation Performance: A Study on Hyperspectral City

The paper investigates how stable band‑selection methods are and how that stability relates to semantic segmentation performance on the Hyperspectral City V2 dataset. Six band‑selection techniques were tested on ten different class‑balanced ROI sets, producing 60 top‑25 band subsets. Results show that Sim‑LP has the highest intra‑method stability, and together with JMIM+CSNR it also delivers the best segmentation results, achieving up to 2.01 mIoU improvement and 18–22× faster CPU inference for a 9‑band subset, though stability does not consistently predict segmentation quality.

By Jiarong Li, Imad Ali Shah, Enda Ward, Martin Glavin, Edward Jones, Brian Deegan