arXiv AI

WasteAssistant: Regulation-Guided Visual Question Answering Framework for Intelligent Waste Segregation and Sustainable Managemen

arXiv:2607. 10610v1 Announce Type: cross Abstract: Efficient waste segregation is critical for sustainable urban management and environmental governance.

arXiv Computer Vision
Aug 31

SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals

SDGBiasBench is a large-scale benchmark suite designed to evaluate and mitigate biases in vision–language models (VLMs) when reasoning about Sustainable Development Goals (SDGs). It contains 500k expert‑involved multiple‑choice questions and 50k regression tasks, allowing assessment of both decision‑level and estimation‑level bias. Experiments show that current VLMs exhibit intrinsic SDG bias, often relying on priors rather than multimodal evidence, and the proposed CADE method significantly reduces this bias, improving accuracy and reducing mean absolute error.

By Zihang Lin, Huaiyuan Qin, Muli Yang, Hongyuan Zhu
arXiv Computer Vision
Aug 25

WADE: A Reasoning-Annotated Benchmark for Multi-Instance Floating-Waste Grounding with Compact Vision-Language Models

arXiv:2608.22950v1 Announce Type: new Abstract: Floating waste in inland waterways threatens aquatic ecosystems and requires timely monitoring under cluttered, multi-object conditions. Existing aquat...

By Md. Asaduzzaman Shuvo, Ahsan Farabi, Md. Abdul Ahad Minhaz, Mahedi Hasan, Israt Khandaker, Ibrahim Khalil Shanto, Muhammad Nomani Kabir
arXiv AI
Sep 18

A Multi-Modal Generative Model for Tomato Disease Leaves Understanding

The paper introduces SOLAR, a multimodal generative model that jointly interprets visual and textual data to understand tomato leaf diseases across six question‑answering tasks. SOLAR aligns visual features with task‑aware language representations using a Fusion Expert module based on a mixture‑of‑experts, enabling it to generate contextually relevant answers for diverse diagnostic tasks. Evaluated on 41,677 images and 216,209 QA pairs, SOLAR outperforms state‑of‑the‑art vision‑only, vision‑language, and task‑specific models in both closed and open‑ended settings, demonstrating superior accuracy, robustness, and multimodal reasoning.

By Khang Nguyen Quoc, Minh-Phuoc Tran, Gia-Han Truong, Luyl-Da Quach
arXiv AI
Jun 10

MMClima: A Framework for Multimodal Climate Science Data and Evaluation

arXiv:2606. 10194v1 Announce Type: cross Abstract: Climate change research increasingly requires AI systems that reason across text, dynamic visual content, and scientific figures, yet existing climate QA benchmarks are small, mostly textual, and cover a narrow range of models.

By Muhammad Umer Sheikh, Hassan Abid, Khawar Shehzad, Ufaq Khan, Muhammad Haris Khan
arXiv Machine Learning
Sep 7

An Integrated Vision-and-Language Pretraining (VLP) and Visual Question Answering (VQA) model to Automate Nondestructive Evaluation Image Analysis

The paper presents ChatNDE Figure to Caption, an AI system that automates the interpretation of nondestructive evaluation (NDE) images. It combines a Vision‑and‑Language Pretraining (VLP) approach using ResNet50 for visual feature extraction and GPT‑2 for natural‑language captioning, evaluated with BLEU scores. Additionally, a Visual Question Answering (VQA) model is integrated to answer specific queries about the images, enhancing interactivity for field inspectors.

By Mehrdad Shafiei Dizaji, Hoda Azari