Robust Promptable Video Object Segmentation
arXiv:2605.12006v2 Announce Type: replace Abstract: The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deploymen...
arXiv:2606. 26734v1 Announce Type: cross Abstract: The impact of real-world noise on Open Vocabulary Object Detectors (OV-ODs) remains poorly understood due to their architectural complexity.
arXiv:2605.12006v2 Announce Type: replace Abstract: The performance of promptable video object segmentation (PVOS) models substantially degrades under input corruptions, which prevents PVOS deploymen...
The paper demonstrates that object detection benchmarks suffer from incomplete annotations, with re-annotation of COCO, Pascal VOC, Cityscapes, and KITTI revealing up to a 60% increase in detected objects, especially small, occluded, or densely packed instances. The authors propose a scalable annotation pipeline that uses multiple annotators per object to capture uncertainty and improve recall, and they introduce two new large-scale benchmarks: an uncertainty-aware detection benchmark and a label error detection benchmark based on real errors. Their findings show that benchmark performance is highly sensitive to annotation quality, yet model rankings remain largely unchanged, highlighting the need for uncertainty-aware evaluation to better reflect real-world ambiguity.
arXiv:2607. 02404v1 Announce Type: cross Abstract: Image encoders trained with LeJEPA can deliver strong features for downstream tasks, but, like other image-level self-supervised methods, typically require large training datasets.
arXiv:2606. 06943v1 Announce Type: cross Abstract: Vision-language models (VLMs) such as CLIP achieve strong zero-shot recognition but remain highly fragile under adversarial perturbations.
arXiv:2602.06652v2 Announce Type: replace Abstract: The robustness of Vision Language Models (VLMs) is commonly assessed through output-level invariance, implicitly assuming that stable predictions r...
arXiv:2609.38010v1 Announce Type: cross Abstract: Modern computer vision models achieve high accuracy when trained on large-scale annotated datasets. In critical domains such as construction safety m...
arXiv:2602.14633v3 Announce Type: replace Abstract: We introduce VIGIL (Visual Inconsistency & Generative In-context Lucidity), a benchmark dataset and framework that provides a fine-grained categori...
arXiv:2602. 18094v2 Announce Type: replace-cross Abstract: Existing Visual-Language Models (VLMs) have achieved significant progress by being trained on massive-scale datasets, typically under the assumption that data are independent and identically distributed (IID).
arXiv:2510. 06596v2 Announce Type: replace-cross Abstract: The performance of machine learning models depends heavily on training data.
arXiv:2605. 07821v2 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection is crucial for ensuring the reliability of deep learning models.
arXiv:2606. 11837v1 Announce Type: cross Abstract: Open-vocabulary scene sketch semantic segmentation aims to assign dense semantic labels to sparse line drawings based on flexible category vocabularies specified at inference time, without relying on pixel-level annotations during training.
arXiv:2608. 05424v1 Announce Type: cross Abstract: Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals.