arXiv:2606. 07882v1 Announce Type: cross Abstract: Different vision neural networks -- trained to classify, contrast, reconstruct, or match images to text -- should have correspondingly different internal representations.
By Yousef Radwan
arXiv:2606. 20364v1 Announce Type: new Abstract: A companion study established a de-biased, cross-model VLM-as-3D-judge that reliably ranks single-image-to-3D mesh quality where cheap geometry and CLIP proxies fall short.
By Ali Asaria, Tony Salomone, Deep Gandhi
arXiv:2308. 04553v4 Announce Type: replace-cross Abstract: Visual recognition models are prone to learning spurious correlations induced by a biased training set where certain conditions $B$ (\eg, Indoors) are over-represented in certain classes $Y$ (\eg, Big Dogs).
By Maan Qraitem, Kate Saenko, Bryan A. Plummer
arXiv:2606. 01973v1 Announce Type: new Abstract: Open-set test-time adaptation (TTA) updates models on new data in the presence of input shifts and unknown output classes.
By Zefeng Li, Evan Shelhamer
arXiv:2603. 24575v2 Announce Type: replace-cross Abstract: Scalable Vector Graphics (SVG) are essential for technical illustration and digital design, offering resolution independence and semantic editability.
By Qijia He, Xunmei Liu, Hammaad Memon, Ziang Li, Zixian Ma, Jaemin Cho, Zhongzheng Ren, Daniel S Weld, Ranjay Krishna
arXiv:2606. 18451v1 Announce Type: new Abstract: Single-image-to-3D generators are improving quickly, but there is no agreed, human-free way to tell whether one generated mesh is better than another.
By Ali Asaria, Tony Salomone, Deep Gandhi
arXiv:2608. 17564v1 Announce Type: cross Abstract: Unified multimodal models (UMMs) are motivated by the hope that understanding and generation reinforce each other but controlled ablations repeatedly find that adding a generation objective leaves understanding flat.
By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong
arXiv:2607. 24440v1 Announce Type: cross Abstract: Vision-language models (VLMs) deployed on consumer hardware must decide when to answer and when to defer, and that decision depends on having a confidence signal that tracks correctness.
By M M Asif Ferdous
arXiv:2606. 06539v1 Announce Type: cross Abstract: Forward-Forward (FF) learning [Hinton, 2022] replaces backpropagation with strictly layer-local goodness updates.
By Yucheng Chen
arXiv:2606. 28401v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have shown strong performance in visual understanding, yet they still suffer from hallucinations, generating content that is not grounded in the image.
By Yunhun Nam, Jongheon Jeong
arXiv:2606. 22437v2 Announce Type: replace-cross Abstract: We conduct a systematic study of 18 widely used vision-language benchmarks and identify three major issues: 1) many items do not rely on visual cues and therefore fail to effectively measure multimodal understanding; 2) many items are already close to performance saturation for current LVLMs, which limits their discriminative power; 3) a small number of anomalous items affect the reliability of evaluation results.
By Wenzhen Yuan, Jiacheng Ruan, Wutao Xiong, Chengping Zhao, Ting Liu, Yuzhuo Fu
arXiv:2605. 30581v2 Announce Type: replace-cross Abstract: Industrial visual sim-to-real is often described as transferring from synthetic images to real images, but industrial deployment usually involves a broader mismatch between available evidence and required decisions.
By Chenxi Tao, Seung-Kyum Choi