The paper introduces FaVOS, a benchmark for Video Object Segmentation (VOS) that focuses on scenarios where target objects appear only intermittently over long videos. It demonstrates that the standard J&F metric can be gamed by empty predictions, allowing trivial models to outperform strong ones like SAM 3. To address this, the authors propose Volumetric J&F, which treats mask sequences as spatio‑temporal volumes, reducing the influence of target‑absence rewards while maintaining sensitivity to segmentation quality and temporal structure.
By Jihwan Hong, Woohyeon Park, Jaeik Kim, Jaeyoung Do
OpenBox is a two‑stage automatic annotation pipeline that uses a 2D vision foundation model to associate 2D image cues with 3D point clouds. It then classifies instances by rigidity and motion state to generate adaptive bounding boxes using class‑specific size statistics, eliminating the need for self‑training. Experiments on Waymo Open Dataset, Lyft Level 5 Perception, and nuScenes show improved accuracy and efficiency over existing baselines.
By In-Jae Lee, Mungyeom Kim, Kwonyoung Ryu, Pierre Musacchio, Jaesik Park
The paper introduces EGGroW, efficient algorithms for computing geodesic Gromov-Wasserstein distances using entropic Sinkhorn-like methods, GenusSink techniques, and random features. It addresses the cubic time complexity of traditional GWD calculations on dense intra-space distance matrices, enabling scalable comparisons of probabilistic distributions on general geodesic manifolds and graph shortest‑path distances. The authors demonstrate EGGroW’s effectiveness in downstream tasks such as 3D pose estimation and partial 3D template recovery, showing accurate results where Euclidean‑based methods fail while maintaining a light computational footprint.
By Krzysztof Marcin Choromanski, Derek Long, Ananya Parashar, Dwaipayan Saha
The paper investigates the gap between retrieval and reading in document vision‑language models, showing that even when the correct page is retrieved, the model often fails to use the text on that page. By comparing answers derived from page images alone versus images plus extracted OCR text, the authors demonstrate that adding OCR text can improve strict accuracy by 13 to 16 points on their FoveDoc‑Bench benchmark. The study also reveals that OCR benefits textual evidence but not charts or figures, and that its advantage diminishes as retrieval quality worsens.
By Qingtao Xia, Siyao Cheng, Jiahua Bao, Jiaxing Du, Jie Liu
The paper introduces a unified framework for multimodal 3D human pose estimation that fuses RGB, LiDAR, and mmWave radar data while incorporating kinematics-based sensor fusion. It presents a black-box subject membership inference attack and a pointwise maximal leakage analysis to assess privacy risks, and proposes a user-level differential privacy method called Action Temporal Stratification to mitigate these risks. The framework is evaluated on the MM-Fi dataset under three experimental protocols, with source code to be released upon acceptance.
By Kaushik Bhargav Sivangi, Fani Deligianni
The paper introduces an automated pipeline that reconstructs editable 3D procedural models of field‑grown maize directly from raw 3D point clouds, eliminating the need for manual tuning or species‑specific training data. It uses a vision‑language model to annotate leaf midlines in rendered views, then applies deterministic geometric algorithms and differentiable NURBS fitting to generate accurate plant descriptors and refine leaf surfaces. The method achieves a median Chamfer distance of 5.4 mm on 100 diverse maize plants and recovers 99.4% of reference leaves with high overlap, outperforming previous semi‑automated approaches.
By Mozhgan Hadadi, Talukder Z. Jubery, Adarsh Krishnamurthy, Baskar Ganapathysubramanian
DEPICT is a new training‑free metric for evaluating text‑to‑image alignment. It replaces fixed reference answers with an agreement rule that compares image‑based and caption‑only responses, weighting questions by how decisively the caption determines them. By merging this agreement score with a holistic score, DEPICT improves negation accuracy dramatically and outperforms existing training‑free metrics while matching or exceeding fine‑tuned evaluators on several benchmarks.
By Vasco Ramos, Sandra Godinho Silva, Joao Magalhaes, Ricardo Rei, Pedro Henrique Martins
arXiv:2602. 13298v4 Announce Type: replace-cross Abstract: This paper presents a controlled comparative study of convolutional neural network (CNN) topology and image classification performance across the architectural families VGG, ResNet, and GoogLeNet, evaluated on CIFAR-10 under a unified training protocol.
By Manfred M. Fischer, Joshua Pitts
ReSCUE is a unified framework for simultaneous sign language translation on unsegmented long‑form videos, aligning training and inference with realistic streaming conditions. It incorporates inference‑aware training to manage partial inputs, non‑signing pauses, and multi‑sentence contexts; stabilized re‑translation for low‑latency, revisable predictions with reduced output flicker; and a sentence commitment mechanism for online segmentation and memory management. Experiments show ReSCUE achieves lower latency and superior translation quality in low‑latency settings, approaching oracle offline systems on long‑form datasets while operating at substantially lower latency.
By Sihan Ren, Gaozheng Li, Yuanshang Quan, Yiming Qin, Fuyi Yang, Chang Liu, Lan Xu, Minye Wu
SCOPE-4D is an endoscopic 4D geometry foundation model that predicts camera parameters, dense geometry, and 3D tissue trajectories from monocular RGB video in a single forward pass. The authors introduce SCOPE-5K, a curated dataset of about 5,000 real and synthetic gastrointestinal endoscopy and laparoscopy clips, and use it for geometric supervised fine‑tuning. Adding Common–Residual Motion (CRM) constraints and trajectory supervision further improves camera and depth estimation and enables dense 3D tissue tracking, as shown by evaluations on public and new benchmarks and a blinded user study.
By Chaoyi Zhou, Zhongpai Gao, Anwesa Choudhuri, Meng Zheng, Benjamin Planche, Run Wang, Terrence Chen, Siyu Huang, Ziyan Wu
The paper presents ViTok, a multi‑teacher distillation approach that combines SigLIP2 and DINOv3‑L to jointly preserve global recognition and dense semantics. By introducing split adaptor heads, asymmetric losses, teacher reweighting, masked image modeling, and PHI‑S feature balancing, the authors achieve higher ImageNet‑1K kNN accuracy and restore ADE20K segmentation performance to match the teacher. The study also reports negative findings, such as limited benefits from scaling to ImageNet22K and interference from additional teachers.
By Hailun Xu, Kanchan Sarkar
arXiv:2610.02597v1 Announce Type: new
Abstract: Vision foundation models such as DINOv2, SigLIP2, and MASt3R develop complementary capabilities from different pretraining objectives, yet their knowle...
By Zhenghao Zhao, Chi Zhang, Qingshuang Chen, Yelin Kim
arXiv:2610.03717v1 Announce Type: cross
Abstract: This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene struc...
By Keerthi Kaashyap, Dennis Anthony, Akshay Krishnan, Nhi Ngoc Nguyen, Jeremy Collins, James Hays, Shreyas Kousik, Animesh Garg
arXiv:2610.03474v1 Announce Type: new
Abstract: Federated learning (FL) enables collaborative training of medical image segmentation models without sharing raw patient data, yet existing approaches a...
By Abhijeet Parida, Zhifan Jiang, Pooneh Roshanitabrizi, Austin Tapp, Maria J. Ledesma-Carbayo, Syed Muhammad Anwar, Ziyue Xu, Marius George Linguraru, Holger R. Roth
arXiv:2610.03248v1 Announce Type: new
Abstract: Panoramic images provide a complete 360-degree field of view, enabling comprehensive scene understanding for embodied perception. However, heterogeneou...
By Pujun Guo, Yuanfan Zheng, Fei Teng, Mengfei Duan, Guoqiang Zhao, Yuheng Zhang, Kai Luo, Kailun Yang
arXiv:2610.03142v1 Announce Type: new
Abstract: Deep neural networks remain vulnerable to adversarial perturbations, which can distort not only predictions but also confidence scores, undermining unc...
By Leo Fillioux, Stergios Christodoulidis, Stergios Christodoulidis, Maria Vakalopoulou, Jose Dolz
The paper introduces CLEAR, a lightweight, task‑agnostic post‑hoc method that enhances evidential robustness in deep learning models without retraining. CLEAR uses held‑out calibration data to map the geometry of the model’s latent space, then generates perturbation views at inference to detect latent conflict. When high conflict is found, CLEAR selectively reduces evidential strength while preserving evidence for latent‑consistent inputs, achieving significant improvements in OOD and adversarial AUROC on ImageNet→CUB and running much faster than competing methods.
By Charmaine Barker, Daniel Bethell, Simos Gerasimou
CrossGMN introduces a graph metanetwork that processes a trained source network and an initialized target network simultaneously, enabling equivariant cross‑architecture weight‑space transformations. By preserving symmetry through cross‑network message passing, CrossGMN can refine target network initializations while remaining invariant to source permutations and equivariant to target permutations. Experiments demonstrate that CrossGMN accelerates knowledge distillation, transfers across datasets without retraining, and unifies compression from diverse source architectures into a common target architecture.
By Adir Dayan, Yam Eitan, Haggai Maron
The paper proposes a structured version of the Information Bottleneck (IB) that separates label-relevant structure from within-condition variation using a dual-bottleneck formulation. It introduces a conditional KL term that targets within-condition information, allowing explicit control over nuisance-like variation in learned representations. Experiments demonstrate improved performance in low-data classification and consistent gains on dense prediction tasks.
By Jingyao Zhang, Yuxuan Li, Lu Han, Ali Anaissi, Nguyen H. Tran
The paper reports a reproducible study of evasion attacks on image and text classifiers. A compact convolutional network on MNIST achieved 98.63% clean accuracy but dropped to 60.20% under FGSM with ε=0.15 and 1.72% with ε=0.30, while PGD reduced accuracy to 32.47% and 0.41%; a bit‑depth‑reduction defense only partially restored performance. In contrast, a DistilBERT model fine‑tuned on the SMS Spam Collection reached 98.75% accuracy and 94.96% F1‑score, yet a sequence of predefined perturbations produced only modest probability shifts and did not flip spam to ham predictions.
By Parker Hummel (Minot State University), Ryne Skabo (Minot State University), Muhammad Abusaqer (Minot State University)