arXiv:2609.24204v1 Announce Type: new
Abstract: Visual segmentation systems encounter objects outside their training distribution during real-world deployment, hindering reliable autonomous systems t...
By Anja Deli\'c, Jurica Runtas, Marin Or\v{s}i\'c, Ivan Markovi\'c, Ivan Petrovi\'c
Visual segmentation systems encounter objects outside their training distribution during real-world deployment, hindering reliable autonomous systems that depend on scene parsing in the perception sta...
Open-world object detection (OWOD) requires a detector to recognize known categories, discover unnamed objects from unseen categories, and incrementally learn newly annotated classes. PROB improves unknown discovery by modeling class-agnostic probabilistic objectness in the decoder-query space.
arXiv:2607. 22212v1 Announce Type: cross Abstract: Visual anomaly detection requires adaptive representations and reliable decision boundaries, particularly when anomalous training samples are scarce and class distributions are highly imbalanced.
By Alireza Dastmalchi Saei, Shervin Rahimzadeh Arashloo
arXiv:2605. 07821v2 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection is crucial for ensuring the reliability of deep learning models.
By Boyang Dai, Chaoqi Chen, Yizhou Yu
Background-Free Objectness Learning (B-FOR) is a dense, class‑agnostic detection framework that learns objectness without treating unlabeled regions as background. It predicts multi‑scale object‑center and scale fields, using spatially structured soft targets to supervise only reliable annotated areas and introduces displacement‑aware scale fields to model object extent. Experiments on PASCAL VOC, MS‑COCO, and Open Images show B‑FOR improves recall by over +10 AR points compared to prior class‑agnostic baselines, with ablation studies confirming the importance of localized supervision and displacement‑aware scaling.
By Dania Batool, Liliana Lo Presti, Marco La Cascia, Filippo Vella
The paper presents the first systematic evaluation of uncertainty quantification (UQ) methods applied to a foundation model for semantic segmentation. By fine‑tuning a lightweight DPT decoder on the pretrained SAM2 encoder, the authors benchmark four UQ approaches—Monte Carlo Dropout, Deep Sub‑Ensemble, Test‑Time Augmentation, and Evidential Deep Learning—across Cityscapes, NYUv2, and two out‑of‑domain settings, comparing segmentation accuracy, calibration, uncertainty quality, and inference time. The results reveal clear trade‑offs between predictive performance, reliability, and computational cost, underscoring both the promise and current limitations of uncertainty‑aware foundation models for real‑world deployment.
By Steven Landgraf, Joceline Hinz, Markus Ulrich
arXiv:2607. 23924v1 Announce Type: cross Abstract: Vision foundation models have enabled strong training-free anomaly detection (AD).
By Jyun-Ze Tang, Po-Han Huang, Ming-Ching Chang, Chih-Fan Hsu, Jeng-Lin Li
arXiv:2607. 23981v1 Announce Type: cross Abstract: Open-world object detection (OWOD) requires a detector to recognize known categories, discover unnamed objects from unseen categories, and incrementally learn newly annotated classes.
By Weijun Tian, Rui Liu
CODE: Cross-Modal Calibration and Dynamic Suppression for Open World Object Detection introduces a unified inference-time framework that addresses semantic ambiguity and over-suppression in multimodal OWOD systems. It comprises Cross-Modal Joint Confidence Calibration, Uncertainty-Guided Universal Objectness Enhancement, and Dynamic Outlier Suppression via Confidence Margin. Experiments on the Real-World Detection benchmark with the OWL‑ViT L/14 backbone show CODE achieving 21.7 U‑mAP and 40.8 K‑mAP, surpassing prior state‑of‑the‑art results by 2.6 and 2.3 points respectively.
By Hao Xu, Zhaoning Shi, Hehe Jin, Bo Ma
The paper identifies a problem in multi‑view anomaly detection called cross‑view information leakage, where fusing multiple inspection views can cause normal features to mask anomalies during reconstruction. To address this, the authors propose GLAD, a framework that uses a Global‑Local Attention Driven approach, combining vision foundation model features with two fusion modules: Multi‑view Merging Attention for local, weighted fusion and Object‑Guided Attention for global context aggregation. Experiments on Real‑IAD and MANTA‑Tiny demonstrate that GLAD outperforms existing methods across various metrics, underscoring the importance of restricting information flow to preserve the reconstruction gap.
By Shang-Fu Chen, Kuan-Chuan Peng, Jhih-Ciang Wu, Wen-Huang Cheng, Kai-Lung Hua
arXiv:2607.09086v2 Announce Type: replace
Abstract: We present Subtoken Vision Transformer (SubViT), a selective image tokenization method for fine-grained visual recognition. Standard Vision Transfo...
By Jie Zhu, Ivy Zhang, Minchul Kim, Xiaoming Liu