arXiv:2604.09531v2 Announce Type: replace-cross
Abstract: Vision-language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition, largely be...
By Guanyu Zhou, Yida Yin, Wenhao Chai, Shengbang Tong, Xingyu Fu, Zhuang Liu
arXiv:2509. 14860v2 Announce Type: replace-cross Abstract: Image classification has traditionally relied on parameter-intensive model training, requiring large-scale annotated datasets and extensive fine tuning to achieve competitive performance.
By Wonduk Seo, Minhyeong Yu, Hyunjin An, Seunghyun Lee
arXiv:2609.37576v1 Announce Type: new
Abstract: With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to captur...
By Yu Zhao, Jiarui Wang, Huiyu Duan, Ye Zhao, Jutao Tang, Juntong Wang, Guangtao Zhai, Xiongkuo Min
arXiv:2606. 26552v1 Announce Type: cross Abstract: The rapid advancement of generative models presents a significant challenge to existing deepfake detection methods, particularly given the widespread dissemination of highly realistic AI-generated images.
By Yangjun Wu, Keyu Yan, Yu Liu, Jingren Zhou, Fei Huang, Rong Zhang, Zhou Zhao, Fei Wu
ReViCo (Real Visual Correction) is a new benchmark that tests Vision Language Models (VLMs) on the task of correcting text errors in real‑world images, requiring deep understanding of visual text and its context. The study evaluates VLMs using both prompt‑based and targeted training approaches, revealing a significant performance gap between current models and humans. The results show that most VLMs struggle to accurately perceive visual text, leading to frequent correction mistakes, thereby underscoring the need for more robust, text‑aware VLMs.
By Bojun Zhang, Junhong Liang, Feifei Zhai, Fengxian Ji, Yu Zhou
arXiv:2607. 14256v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly deployed for nuanced content safety and moderation tasks, yet they remain vulnerable to adversarial attacks and out-of-distribution edge cases.
By Genglin Liu, Muye Zhang, Krishnamurthy Viswanathan, Nichole J. Hansen, Bla\v{z} Bratani\v{c}, Nathan L Clement, Shalini Ghosh, Ariel Fuxman