arXiv:2607. 16243v1 Announce Type: cross Abstract: Multimodal industrial anomaly inspection assistants are a critical component of next-generation smart factories, enabling interactive vision-language-based querying.
By Anushiya Arunan, Xin Li, Yan Qin, U-Xuan Tan, Nhu Khue Vuong, Xiaoli Li, Chau Yuen
The paper introduces VertiCue-Bench, a diagnostic benchmark designed to test whether multimodal large language models (MLLMs) can perceive, ground, and utilize vertical structure information in remote-sensing natural scenes. It presents a three-stage framework—Perception, Grounding, Utilization—and a Representation Intervention Spectrum across various modalities to evaluate ten state-of-the-art models. The study finds a significant Vertical Structure Utilization Gap: while models show some geometric perception, they struggle to accurately link vertical evidence to spatial entities and incorporate it into semantic decisions.
By Jing Huang, Duanchu Wang, Junjie Yang, Zihang Cheng, Cheng Li, Lin Cui, Zhouyi Wu, Di Wang
arXiv:2606. 17904v1 Announce Type: new Abstract: Language models increasingly serve as advisory systems in maintenance operations.
By Guillermo Gil de Avalle, Laura Maruster, Shaina Raza, Christos Emmanouilidis
arXiv:2607. 24539v1 Announce Type: new Abstract: Multimodal large language models (LLMs) can combine topology, measurements, and incident text for grid diagnosis, yet answer accuracy does not establish that task-appropriate evidence was used.
By Tianqiao Zhao, Meng Yue, Jianhui Wang
arXiv:2607. 09544v1 Announce Type: cross Abstract: Despite strong performance on many multimodal tasks, vision-language models (VLMs) still struggle with basic object counting.
By Ahmed Oumar El-Shangiti, Abzal Nurgazy, Hilal AlQuabeh, Nikolai Rozanov, Kentaro Inui
Existing video benchmarks evaluate action recognition on consumer videos, egocentric recordings, or simulated industrial environments. They do not test vision-language models under the visual and procedural conditions of real industrial CCTV, where workers appear as distant figures amid dust, steam, low light, glare, occlusion, and overlapping activities.