arXiv:2607.05568v2 Announce Type: replace-cross
Abstract: Compact primitive abstractions represent 3D shapes with a few geometric primitives while preserving recognizable components. Learned methods...
By Gregor Kobsik, Tim Elsner, Leif Kobbelt
arXiv:2609.38391v1 Announce Type: new
Abstract: Document text forgery has evolved beyond simple pixel-level manipulation: modern attacks alter not only the appearance of a document but also its meani...
By Kirill Koltsov, Aleksandr Gushchin, Dmitriy Vatolin, Anastasia Antsiferova
arXiv:2609.38476v1 Announce Type: new
Abstract: Synthetic data are most valuable where general-purpose datasets cannot provide the domain-specific priors a task requires, and where manual annotation...
By Saptarshi Neil Sinha, Paul Julius K\"uhn, Michael Weinmann
arXiv:2609.38714v1 Announce Type: new
Abstract: We describe our winning entry to the Waymo Open Dataset 2D Video Panoptic Segmentation Challenge. The task asks for a semantic class at every pixel of...
By Jinghan Yang
arXiv:2609.39266v1 Announce Type: new
Abstract: Fine-grained vision-language alignment in chest radiography enables zero-shot classification, grounding, and segmentation without task-specific annotat...
By Qixing Zhao, Jinpeng Li
arXiv:2609.40244v1 Announce Type: new
Abstract: Mobile robots and vehicles carry synchronized multi-camera rigs, yet many streaming 3D foundation models are designed for monocular input, leaving effi...
By Yufei Wei, Shuhao Ye, Qi Wang, Xin Zheng, Qing Huang, Rong Xiong, Yue Wang
arXiv:2609.40347v1 Announce Type: new
Abstract: We introduce VideoMSN, a Masked Siamese Network framework for efficient self-supervised spatio-temporal representation learning in videos. Instead of r...
By Owais Iqbal, Sudipta Sarkar, Shyam Marjit, Omprakash Chakraborty, Anirban Chakraborty, Abir Das
arXiv:2601.08355v3 Announce Type: replace
Abstract: Vision-Language Models (VLMs) are increasingly deployed in autonomous driving and embodied AI systems, where reliable perception is critical for sa...
By Guo Cheng, Huang Li
arXiv:2603.18792v3 Announce Type: replace
Abstract: Uncertainty quantification (UQ) is crucial in safety-critical applications such as medical image segmentation. Total uncertainty is typically decom...
By Jakob L{\o}nborg Christensen, Vedrana Andersen Dahl, Morten Rieger Hannemose, Anders Bjorholm Dahl, Christian F. Baumgartner
arXiv:2603.24528v2 Announce Type: replace
Abstract: Vision-language models like CLIP are trained with the objective of aligning text and image pairs. Beyond text prompts alone, recent works show that...
By Dipam Goswami, Simone Magistri, Gido M. van de Ven, Bart{\l}omiej Twardowski, Andrew D. Bagdanov, Tinne Tuytelaars, Joost van de Weijer
arXiv:2607.03470v2 Announce Type: replace
Abstract: Synthesizing physically accurate mirror reflections remains a fundamental challenge for modern text-to-image diffusion models, which are increasing...
By Xuan-Bach Mai, Duy-Phuc Nguyen, Quoc-Van Le, Tam V. Nguyen, Thanh-Toan Do, Huu Le, Duong-Van Nguyen, Minh-Triet Tran, Trung-Nghia Le
arXiv:2609.34863v2 Announce Type: replace
Abstract: Multimodal large language models (MLLMs) have approached image segmentation by reasoning about visual content and predicting target locations. Thei...
By Cilin Yan, Yilun Qiu, Wanyang Zhang, Rui Zu, Xiaolong Jiang, Jiayin Cai, Yao Hu
arXiv:2602.10124v2 Announce Type: replace-cross
Abstract: Cycling is reported by an average of 35% of adults at least once per week across 28 countries, and as vulnerable road users directly exposed...
By Haining Ding, Chenxi Wang, Simon Ladouce, Michal Gath-Morad
arXiv:2511.11286v4 Announce Type: replace-cross
Abstract: Out-of-domain (OOD) robustness is challenging to achieve in real-world computer vision, especially in unsupervised domain adaptation scenario...
By Ruoqi Wang, Haitao Wang, Shaojie Guo, Qiong Luo
arXiv:2609.38282v1 Announce Type: new
Abstract: Vision-language models may rewrite anomalous text in images into linguistically plausible expressions, compromising OCR transcription faithfulness. Seq...
By Baode Wang, Zuming Huang, Kexuan Ren, Jun Huang, Wei Chu
arXiv:2609.39182v1 Announce Type: cross
Abstract: World Models are appearing as the next major frontier in computer vision. However, their robustness is currently largely unexplored. We identify the...
By Ali J Alrasheed, Aryan Yazdan Parast, Basim Azam, James Bailey, Naveed Akhtar
arXiv:2609.39453v1 Announce Type: cross
Abstract: Speech emotion recognition (SER) is the task of assigning emotion labels to utterances. Early systems relied on acoustic features, whereas recent app...
By Hezhao Zhang, Thomas Hain
arXiv:2609.39704v1 Announce Type: cross
Abstract: This paper demonstrate that whether masking-based token pruning helps or hurts worst-group robustness can be predicted before deployment, without lab...
By Muhammad Zawish, Steven Davy
arXiv:2609.38856v1 Announce Type: new
Abstract: The equirectangular projection (ERP) is widely used for panoramic depth estimation, but its spatially varying distortion makes geometry-consistent feat...
By Zhijie Shen, Chunyu Lin, Shuai Zheng, Feng Li, Runmin Cong, Huihui Bai, Yao Zhao
arXiv:2609.39083v1 Announce Type: new
Abstract: Super-resolution and quality enhancement of 1.5\,T brain MRI are normally validated with image-fidelity metrics, although their purpose is to improve d...
By Kavitha Viswanathan, Harsh Choudhary, Amit Sethi