arXiv:2512. 09062v2 Announce Type: replace-cross Abstract: Accurate 3D scene interpretation in active construction sites is essential for progress monitoring, safety assessment, and digital twin development.
By Seongyong Kim, Yong Kwon Cho
arXiv:2608. 07651v1 Announce Type: new Abstract: Large language models (LLMs) show promise in medical image interpretation but suffer from hallucination, limited accuracy, and run-to-run inconsistency.
By Jalil Jalili, Hossein Taghizad, Anuwat Jiravarnsirikul, Christopher Bowd, Akram Belghith, Raheleh Kafieh, Christopher A. Girkin, Sally L. Baxter, Robert N. Weinreb, Linda M. Zangwill, Mark Christopher
arXiv:2608. 08935v1 Announce Type: new Abstract: This work presents a unified multimodal AI system for damage assessment that integrates retrieval-augmented generation (RAG) models, thermal spectrum perception, vision foundation model pipelines, and exploratory wireless signal sensing.
By Kalelo Dukuray, Israel Pina, Evan Perez, Erika Ardiles-Cruz, Jie Wei
arXiv:2608. 07917v1 Announce Type: new Abstract: Chinese historical documents preserve valuable cultural heritage, but many collections remain accessible only as scanned page images, preventing full-text retrieval, collation, and computational analysis.
By Zhongheng Zhou, Yi Sun, Huiguo He, Yuyi Zhang, Peirong Zhang, Yulin Fang, Dezhi Peng, Minghui Liao, Lianwen Jin
arXiv:2608. 09091v1 Announce Type: cross Abstract: Transfer learning is particularly useful in settings with limited training data, and within image classification it is common to transfer learn upon massive datasets like ImageNet , CIFAR-100, or COCO .
By Jing Ning, James D. Braza
arXiv:2608. 09846v1 Announce Type: new Abstract: This paper presents a methodological framework for real-time climate risk assessment using data-driven nowcasting techniques to enhance supply chain resilience in Colombian agricultural contexts.
By Hernan J. Silva-Sosa
arXiv:2608. 07886v1 Announce Type: cross Abstract: Vision-language grounding connects language to visual content, yet most existing formulations reduce grounding to a unidirectional localization problem: given a prespecified text phrase or category name, identify the corresponding image region.
By Jieyu Zhang, Ziqi Gao, Luke Zettlemoyer, Ranjay Krishna
arXiv:2608. 09236v1 Announce Type: new Abstract: Federated learning enables privacy-preserving collaboration across distributed devices without centralizing local data.
By Jaeheon Kim, Hokeun Kim, Bong Jun Choi
arXiv:2608. 07575v1 Announce Type: cross Abstract: Confocal microscopy of optically cleared and swelled tissue resolves complex biological structures in 3D, but such acquisitions are highly anisotropic: along the under-sampled axial direction the structure can appear discontinuous, hampering reconstruction and automated quantitative analysis.
By Arash Fatehi, Robin Ebbestad, Linus Butt, Hans Blom, Sigrid Lundberg, Hannes Olauson, Hjalmar Brismar, David Unnersj\"o-Jess, Thomas Benzing, Katarzyna Bozek
arXiv:2608. 00442v2 Announce Type: replace-cross Abstract: Medical anomaly detection identifies abnormal images and localizes lesions under scarce supervision while generalizing across organs and modalities.
By Yibo Wan, Jinyu Cai, See-kiong Ng
arXiv:2608. 08873v1 Announce Type: cross Abstract: Facial Emotion Recognition (FER) is an important task that has significant implications across various fields such as biometrics, health, and human-computer interaction.
By Aya Manel Zitouni, Aicha Zenakhri, Karim Haroun, Larbi Boubchir
arXiv:2608. 09101v1 Announce Type: cross Abstract: Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misaligned; high overlap scores may then reflect agreement with imperfect labels rather than faithfulness to the image, creating an evaluation paradox.
By Shuaishuai Cao, Shuwei Peng, Meng Tang, Min Huang, Youjin Wang, Jie Chen, Jing Ouyang, Zhiwei Zhai
arXiv:2608. 09928v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control.
By Hunar Batra, Lachin Naghashyar, Ashkan Khakzar, Philip Torr, Christian Schroeder de Witt, Constantin Venhoff, Ronald Clark
arXiv:2608. 08961v1 Announce Type: new Abstract: AI training's rising resource intensity is straining electricity supplies and carbon budgets, motivating systematic study of memory-efficient training on constrained hardware.
By Sarthak Mahapatra, Zihan Zhou, Khatoon Khedri, Mehdi Hosseinzadeh, Reza Rawassizadeh
arXiv:2601. 03100v3 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) typically rely on a single late-layer feature from a frozen vision encoder, leaving the encoder's rich hierarchy of visual cues under-utilized.
By Chenchen Lin, Sanbao Su, Rachel Luo, Yuxiao Chen, Yan Wang, Marco Pavone, Fei Miao
arXiv:2608. 09818v1 Announce Type: cross Abstract: Reliable medical image understanding requires models to connect clinical language and visual reasoning with pixel-level grounding.
By Haoyu Yang, Meixing Shi, Zengjie Chen, Haoran Sun, Haitao Leng, Xiaoming Shi, Yuxiang Cai, Yankai Jiang
arXiv:2608. 09360v1 Announce Type: cross Abstract: The demand for maritime surveillance has given rise to the need for monitoring fishing vessel activities, particularly in addressing the challenge of "dark vessels" that operate without Automatic Identification System (AIS) transmission.
By Shantakar Mohanty, Prasun Kumar Gupta, Raian Vargas Maretto
arXiv:2608. 07577v1 Announce Type: cross Abstract: A closed-set detector for autonomous driving must assign every object one of a fixed set of labels.
By Felix Schaller
arXiv:2608. 08308v1 Announce Type: cross Abstract: Modern vision systems must operate in "open-world" settings, where models must recognize known categories and detect unseen or anomalous content.
By Anastasios Romanos Varvarigos, Nikos Giakoumoglou, Tania Stathaki
arXiv:2608. 07750v1 Announce Type: cross Abstract: Deep Neural Networks (DNNs) have found successful deployment in numerous vision perception systems.
By Cong Chen, Jean-Philippe Monteuuis, Jonathan Petit