arXiv:2607. 02551v1 Announce Type: cross Abstract: Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception.
By Yankai Yang, Yancheng Long, Bin Wen, Fan Yang, Tingting Gao, Han Li, Shuo Yang
arXiv:2606. 28724v1 Announce Type: cross Abstract: Understanding and localizing subtle changes between paired images is critical for tasks such as surveillance and image editing.
By Jinhong Hu, Xiaoping Wang, Shuyin Huang, Guojin Zhong, Kaitai Liu, Kai Lu
arXiv:2605. 15375v2 Announce Type: replace-cross Abstract: Remote sensing change detection (RSCD) localises changes between two images of the same geographic region.
By Bla\v{z} Rolih, Matic Fu\v{c}ka, Filip Wolf, Luka \v{C}ehovin Zajc
arXiv:2607. 21318v1 Announce Type: cross Abstract: Replacing an object with one that differs in category or shape requires complete source removal, natural target formation unconstrained by the source silhouette, and preservation of unrelated content.
By Jian Zhang, Zhijun Zhang
arXiv:2606. 04433v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used in multi-image, multi-turn agentic settings where decisions depend on visual changes.
By Zirui Wang, Junwei Yu, Adam Yala, David M. Chan, Joseph E. Gonzalez, Trevor Darrell
arXiv:2606. 00987v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have shown strong visual understanding and language-guided grounding abilities, yet their capacity for multi-temporal visual reasoning remains underexplored.
By Bingyu Li, Da Zhang, Tao Huo, Zhiyuan Zhao, Junyu Gao, Xuelong Li