Computer vision

Detection, segmentation, depth and recognition research, plus the vision backbones that keep displacing the last generation.

3,374 stories · RSS feed

arXiv Computer Vision
5d ago

When Predicting Nothing Beats SAM 3: Revisiting Evaluation in Video Object Segmentation

The paper introduces FaVOS, a benchmark for Video Object Segmentation (VOS) that focuses on scenarios where target objects appear only intermittently over long videos. It demonstrates that the standard J&F metric can be gamed by empty predictions, allowing trivial models to outperform strong ones like SAM 3. To address this, the authors propose Volumetric J&F, which treats mask sequences as spatio‑temporal volumes, reducing the influence of target‑absence rewards while maintaining sensitivity to segmentation quality and temporal structure.

By Jihwan Hong, Woohyeon Park, Jaeik Kim, Jaeyoung Do
arXiv Computer Vision
5d ago

OpenBox: Annotate Any Bounding Boxes in 3D

OpenBox is a two‑stage automatic annotation pipeline that uses a 2D vision foundation model to associate 2D image cues with 3D point clouds. It then classifies instances by rigidity and motion state to generate adaptive bounding boxes using class‑specific size statistics, eliminating the need for self‑training. Experiments on Waymo Open Dataset, Lyft Level 5 Perception, and nuScenes show improved accuracy and efficiency over existing baselines.

By In-Jae Lee, Mungyeom Kim, Kwonyoung Ryu, Pierre Musacchio, Jaesik Park
arXiv Computer Vision
5d ago

Unlocking Geodesic Gromov-Wasserstein Distances for 3D Modeling

The paper introduces EGGroW, efficient algorithms for computing geodesic Gromov-Wasserstein distances using entropic Sinkhorn-like methods, GenusSink techniques, and random features. It addresses the cubic time complexity of traditional GWD calculations on dense intra-space distance matrices, enabling scalable comparisons of probabilistic distributions on general geodesic manifolds and graph shortest‑path distances. The authors demonstrate EGGroW’s effectiveness in downstream tasks such as 3D pose estimation and partial 3D template recovery, showing accurate results where Euclidean‑based methods fail while maintaining a light computational footprint.

By Krzysztof Marcin Choromanski, Derek Long, Ananya Parashar, Dwaipayan Saha
arXiv Computer Vision
5d ago

Found but Not Read: When Extracted Text Closes the Retrieval-Reading Gap in Document Vision-Language Models

The paper investigates the gap between retrieval and reading in document vision‑language models, showing that even when the correct page is retrieved, the model often fails to use the text on that page. By comparing answers derived from page images alone versus images plus extracted OCR text, the authors demonstrate that adding OCR text can improve strict accuracy by 13 to 16 points on their FoveDoc‑Bench benchmark. The study also reveals that OCR benefits textual evidence but not charts or figures, and that its advantage diminishes as retrieval quality worsens.

By Qingtao Xia, Siyao Cheng, Jiahua Bao, Jiaxing Du, Jie Liu
arXiv Computer Vision
5d ago

Kinematics-Induced Multimodal 3D Human Pose Estimation with Subject-Level Privacy

The paper introduces a unified framework for multimodal 3D human pose estimation that fuses RGB, LiDAR, and mmWave radar data while incorporating kinematics-based sensor fusion. It presents a black-box subject membership inference attack and a pointwise maximal leakage analysis to assess privacy risks, and proposes a user-level differential privacy method called Action Temporal Stratification to mitigate these risks. The framework is evaluated on the MM-Fi dataset under three experimental protocols, with source code to be released upon acceptance.

By Kaushik Bhargav Sivangi, Fani Deligianni
arXiv Computer Vision
5d ago

A Vision-Language Model (VLM)-based Pipeline for End-to-End Procedural Modeling of Field-Grown Maize from Point Clouds

The paper introduces an automated pipeline that reconstructs editable 3D procedural models of field‑grown maize directly from raw 3D point clouds, eliminating the need for manual tuning or species‑specific training data. It uses a vision‑language model to annotate leaf midlines in rendered views, then applies deterministic geometric algorithms and differentiable NURBS fitting to generate accurate plant descriptors and refine leaf surfaces. The method achieves a median Chamfer distance of 5.4 mm on 100 diverse maize plants and recovers 99.4% of reference leaves with high overlap, outperforming previous semi‑automated approaches.

By Mozhgan Hadadi, Talukder Z. Jubery, Adarsh Krishnamurthy, Baskar Ganapathysubramanian
arXiv Computer Vision
5d ago

DEPICT: Scoring Text-to-Image Alignment by Answer Agreement

DEPICT is a new training‑free metric for evaluating text‑to‑image alignment. It replaces fixed reference answers with an agreement rule that compares image‑based and caption‑only responses, weighting questions by how decisively the caption determines them. By merging this agreement score with a holistic score, DEPICT improves negation accuracy dramatically and outperforms existing training‑free metrics while matching or exceeding fine‑tuned evaluators on several benchmarks.

By Vasco Ramos, Sandra Godinho Silva, Joao Magalhaes, Ricardo Rei, Pedro Henrique Martins
arXiv Computation and Language
5d ago

ReSCUE: Re-translation with Sentence Commitment for Unsegmented Long-Form Simultaneous Sign Language Translation

ReSCUE is a unified framework for simultaneous sign language translation on unsegmented long‑form videos, aligning training and inference with realistic streaming conditions. It incorporates inference‑aware training to manage partial inputs, non‑signing pauses, and multi‑sentence contexts; stabilized re‑translation for low‑latency, revisable predictions with reduced output flicker; and a sentence commitment mechanism for online segmentation and memory management. Experiments show ReSCUE achieves lower latency and superior translation quality in low‑latency settings, approaching oracle offline systems on long‑form datasets while operating at substantially lower latency.

By Sihan Ren, Gaozheng Li, Yuanshang Quan, Yiming Qin, Fuyi Yang, Chang Liu, Lan Xu, Minye Wu
arXiv Computer Vision
5d ago

SCOPE-4D: Endoscopic 4D Geometry Foundation Models

SCOPE-4D is an endoscopic 4D geometry foundation model that predicts camera parameters, dense geometry, and 3D tissue trajectories from monocular RGB video in a single forward pass. The authors introduce SCOPE-5K, a curated dataset of about 5,000 real and synthetic gastrointestinal endoscopy and laparoscopy clips, and use it for geometric supervised fine‑tuning. Adding Common–Residual Motion (CRM) constraints and trajectory supervision further improves camera and depth estimation and enables dense 3D tissue tracking, as shown by evaluations on public and new benchmarks and a blinded user study.

By Chaoyi Zhou, Zhongpai Gao, Anwesa Choudhuri, Meng Zheng, Benjamin Planche, Run Wang, Terrence Chen, Siyu Huang, Ziyan Wu
arXiv Computer Vision
5d ago

ViTok: Improving Dense Semantics in AM-RADIO-Style Multi-Teacher Distillation with PHI-S and Masked Image Modelling

The paper presents ViTok, a multi‑teacher distillation approach that combines SigLIP2 and DINOv3‑L to jointly preserve global recognition and dense semantics. By introducing split adaptor heads, asymmetric losses, teacher reweighting, masked image modeling, and PHI‑S feature balancing, the authors achieve higher ImageNet‑1K kNN accuracy and restore ADE20K segmentation performance to match the teacher. The study also reports negative findings, such as limited benefits from scaling to ImageNet22K and interference from additional teachers.

By Hailun Xu, Kanchan Sarkar
arXiv Computer Vision
5d ago

Fed-ADApt: Federated Anytime Depth Adaptation for Resource-Aware Medical Image Segmentation

arXiv:2610.03474v1 Announce Type: new Abstract: Federated learning (FL) enables collaborative training of medical image segmentation models without sharing raw patient data, yet existing approaches a...

By Abhijeet Parida, Zhifan Jiang, Pooneh Roshanitabrizi, Austin Tapp, Maria J. Ledesma-Carbayo, Syed Muhammad Anwar, Ziyue Xu, Marius George Linguraru, Holger R. Roth
arXiv Machine Learning
Oct 3

Robust Evidential Learning Through Latent Consistency

The paper introduces CLEAR, a lightweight, task‑agnostic post‑hoc method that enhances evidential robustness in deep learning models without retraining. CLEAR uses held‑out calibration data to map the geometry of the model’s latent space, then generates perturbation views at inference to detect latent conflict. When high conflict is found, CLEAR selectively reduces evidential strength while preserving evidence for latent‑consistent inputs, achieving significant improvements in OOD and adversarial AUROC on ImageNet→CUB and running much faster than competing methods.

By Charmaine Barker, Daniel Bethell, Simos Gerasimou
arXiv Machine Learning
Oct 3

CrossGMN: Graph Metanetworks for Cross-Architecture Weight-Space Transformations

CrossGMN introduces a graph metanetwork that processes a trained source network and an initialized target network simultaneously, enabling equivariant cross‑architecture weight‑space transformations. By preserving symmetry through cross‑network message passing, CrossGMN can refine target network initializations while remaining invariant to source permutations and equivariant to target permutations. Experiments demonstrate that CrossGMN accelerates knowledge distillation, transfers across datasets without retraining, and unifies compression from diverse source architectures into a common target architecture.

By Adir Dayan, Yam Eitan, Haggai Maron
arXiv Machine Learning
Oct 3

Rethinking the Information Bottleneck: Structured Decomposition under Label-Induced Partitions

The paper proposes a structured version of the Information Bottleneck (IB) that separates label-relevant structure from within-condition variation using a dual-bottleneck formulation. It introduces a conditional KL term that targets within-condition information, allowing explicit control over nuisance-like variation in learned representations. Experiments demonstrate improved performance in low-data classification and consistent gains on dense prediction tasks.

By Jingyao Zhang, Yuxuan Li, Lu Han, Ali Anaissi, Nguyen H. Tran
arXiv Machine Learning
Oct 3

Evasion Attacks: How Adversarial Noise Bypasses ML Classifiers

The paper reports a reproducible study of evasion attacks on image and text classifiers. A compact convolutional network on MNIST achieved 98.63% clean accuracy but dropped to 60.20% under FGSM with ε=0.15 and 1.72% with ε=0.30, while PGD reduced accuracy to 32.47% and 0.41%; a bit‑depth‑reduction defense only partially restored performance. In contrast, a DistilBERT model fine‑tuned on the SMS Spam Collection reached 98.75% accuracy and 94.96% F1‑score, yet a sequence of predefined perturbations produced only modest probability shifts and did not flip spam to ham predictions.

By Parker Hummel (Minot State University), Ryne Skabo (Minot State University), Muhammad Abusaqer (Minot State University)