SDGBiasBench is a large-scale benchmark suite designed to evaluate and mitigate biases in vision–language models (VLMs) when reasoning about Sustainable Development Goals (SDGs). It contains 500k expert‑involved multiple‑choice questions and 50k regression tasks, allowing assessment of both decision‑level and estimation‑level bias. Experiments show that current VLMs exhibit intrinsic SDG bias, often relying on priors rather than multimodal evidence, and the proposed CADE method significantly reduces this bias, improving accuracy and reducing mean absolute error.
By Zihang Lin, Huaiyuan Qin, Muli Yang, Hongyuan Zhu
arXiv:2508. 17117v3 Announce Type: replace-cross Abstract: Existing plant-disease datasets target classification and detection, leaving vision-language models unable to support interactive, reasoning-based diagnosis.
By Syed Nazmus Sakib, Nafiul Haque, Mohammad Zabed Hossain, Shifat E. Arman
arXiv:2608.22950v1 Announce Type: new
Abstract: Floating waste in inland waterways threatens aquatic ecosystems and requires timely monitoring under cluttered, multi-object conditions. Existing aquat...
By Md. Asaduzzaman Shuvo, Ahsan Farabi, Md. Abdul Ahad Minhaz, Mahedi Hasan, Israt Khandaker, Ibrahim Khalil Shanto, Muhammad Nomani Kabir
Floating waste in inland waterways threatens aquatic ecosystems and requires timely monitoring under cluttered, multi-object conditions. Existing aquatic-waste datasets provide limited geographic cove...
Artificial intelligence for plant disease analysis has advanced from task-specific classifiers to multi-modal models capable of jointly interpreting visual and textual information. However, practical...
The paper introduces SOLAR, a multimodal generative model that jointly interprets visual and textual data to understand tomato leaf diseases across six question‑answering tasks. SOLAR aligns visual features with task‑aware language representations using a Fusion Expert module based on a mixture‑of‑experts, enabling it to generate contextually relevant answers for diverse diagnostic tasks. Evaluated on 41,677 images and 216,209 QA pairs, SOLAR outperforms state‑of‑the‑art vision‑only, vision‑language, and task‑specific models in both closed and open‑ended settings, demonstrating superior accuracy, robustness, and multimodal reasoning.
By Khang Nguyen Quoc, Minh-Phuoc Tran, Gia-Han Truong, Luyl-Da Quach
arXiv:2606. 10194v1 Announce Type: cross Abstract: Climate change research increasingly requires AI systems that reason across text, dynamic visual content, and scientific figures, yet existing climate QA benchmarks are small, mostly textual, and cover a narrow range of models.
By Muhammad Umer Sheikh, Hassan Abid, Khawar Shehzad, Ufaq Khan, Muhammad Haris Khan
arXiv:2604.12335v2 Announce Type: replace-cross
Abstract: Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as...
By Tanzila Rahman, Renjie Liao, Leonid Sigal
arXiv:2608. 01664v1 Announce Type: cross Abstract: We present our ImageCLEF 2026 Multimodal Reasoning system for the Visual Multiple Choice Question Answering (Visual MCQ) and Visual Open Question Answering (Visual OpenQA) subtasks.
By Mohamed Basem, Vincent Christlein
arXiv:2511.20022v3 Announce Type: replace-cross
Abstract: Recent advancements in multimodal large language models (MLLMs) have shown strong understanding of driving scenes, drawing interest in their...
By Seungjun Yu, Seonho Lee, Namho Kim, Jaeyo Shin, Junsung Park, Wonjeong Ryu, Raehyuk Jung, Hyunjung Shim
arXiv:2610.01180v1 Announce Type: new
Abstract: Despite the strong performance of Vision-Language Models (VLMs) on a wide range of visual question answering (VQA) tasks, these models consistently str...
By Yuliang Cai, Mohammad Rostami, Jesse Thomason
The paper presents ChatNDE Figure to Caption, an AI system that automates the interpretation of nondestructive evaluation (NDE) images. It combines a Vision‑and‑Language Pretraining (VLP) approach using ResNet50 for visual feature extraction and GPT‑2 for natural‑language captioning, evaluated with BLEU scores. Additionally, a Visual Question Answering (VQA) model is integrated to answer specific queries about the images, enhancing interactivity for field inspectors.
By Mehrdad Shafiei Dizaji, Hoda Azari