arXiv:2609.06245v1 Announce Type: cross
Abstract: Multimodal Large Language Models (MLLMs) perform strongly on general visual understanding tasks such as visual question answering, yet they often str...
By Yixin Wan, Tianle Zheng, Kai-Wei Chang
arXiv:2608. 07435v1 Announce Type: cross Abstract: Vision-language models (VLMs) are improving rapidly, but benchmark development lags behind, making weaknesses hard to identify.
By Zixuan Lan, Luzhe Sun, Matthew R. Walter, Jiawei Zhou
Vision-language models (VLMs) have shown strong capabilities in generating visualization code from textual or visual specifications. However, real-world visualization authoring is inherently iterative: users frequently revise existing visualizations to repair flawed charts or adapt them to desired styles.
The paper introduces the Semantic Drift Protocol (SDP), a Telephone Game-inspired method to evaluate how well unified models preserve meaning when alternating between image-to-text (I2T) and text-to-image (T2I) generation over multiple generations. It defines Mean Cumulative Drift (MCD) and Multi-Generation GenEval (MGG) as metrics for semantic retention, and presents a new benchmark of 400 image‑text pairs from NoCaps and DOCCI to stress-test models beyond COCO. Applying SDP to seven models shows that models with strong single‑pass scores can still suffer severe semantic drift, revealing catastrophic failure modes that isolated benchmarks miss.
By Sabbir Mollah, Rohit Gupta, Sirnam Swetha, Qingyang Liu, Ahnaf Munir, Mubarak Shah
arXiv:2607. 28967v1 Announce Type: cross Abstract: Prompt tuning adapts vision--language models with few trainable parameters, but existing approaches trade off efficiency and adaptation: static textual prompts can overfit source classes, image-conditioned prompts add per-instance computation, and multimodal tuning modifies the visual branch.
By Pouya Parsa, Raoof Zare Moayedi, Seongjin Choi
The paper proposes a weakly supervised remote sensing change detection method that uses change captions as the sole supervision signal, eliminating the need for pixel‑level change masks. It introduces a caption‑driven generation pipeline to create bi‑temporal image pairs with controlled changes and a Semantic‑Appearance Agreement Framework (SAAF) that fuses caption‑grounded semantic responses with RGB differences for accurate change localization. Experiments on the Flair‑RSGen and WHU‑CDC datasets demonstrate that SAAF outperforms existing limited‑supervision baselines in macro‑averaged IoU and F1 metrics.
By Yuan Qian, Jie Ma
Large language models (LLMs) have demonstrated growing competence in web page generation. However, existing text-driven approaches rely on complex prompts that impose substantial demands on users and offer limited expressivity for page layout and cross-page visual coherence.
arXiv:2607. 22705v1 Announce Type: cross Abstract: Object-centric learning aims to represent scenes as objects whose properties can be reused in new combinations.
By Anuraag Gadehothur Karnam, Tarunesh Sathish
arXiv:2607. 06306v1 Announce Type: cross Abstract: Large language models (LLMs) have demonstrated growing competence in web page generation.
By Grace Man Chen, Litao Guo, Yifan Wu, Yiyu Chen, Yenchi Tseng, Sicheng Liu, Yuyu Luo, Ying-Cong Chen
arXiv:2609.10356v1 Announce Type: new
Abstract: Long-term change understanding from images of the same place revisited over time is a challenging task with applications in map maintenance and urban i...
By Benedetta Liberatori, Nermin Samet, Paolo Rota, Matthieu Cord, Elisa Ricci, Andrei Bursuc, Monika Wysocza\'nska
The paper introduces visual adaptations of counterfactual tests—vCT and vCCT—to evaluate whether chain-of-thought explanations in vision‑language models faithfully reflect the visual evidence driving predictions. Using these tests, the authors benchmark eight open‑source VLMs on two datasets and find that CoTs often fail to track visual evidence, sometimes omitting removed objects or mentioning them inconsistently. They also release two new datasets, Counter‑SNLI‑VE and Counter‑A‑OKVQA, consisting of image pairs that differ by a single object to facilitate further research.
By Bayar Menzat, Maximilian S\"uss, Ruizhi Wang, Benno Steinegger, Thomas Lukasiewicz, Oana-Maria Camburu
arXiv:2602. 19946v5 Announce Type: replace-cross Abstract: Recent text-to-image (T2I) diffusion models produce visually stunning images and demonstrate excellent prompt following.
By Krzysztof Adamkiewicz, Brian Bernhard Moser, Stanislav Frolov, Tobias Christian Nauen, Federico Raue, Andreas Dengel