arXiv AI By Khush Kataruka, Harshit Maurya, Anuja Vats, Murari Mandal, Kiran Raja, Praveen Kumar Chandaliya

WasteAssistant: Regulation-Guided Visual Question Answering Framework for Intelligent Waste Segregation and Sustainable Managemen

Read the original on arXiv AI →

arXiv:2607. 10610v1 Announce Type: cross Abstract: Efficient waste segregation is critical for sustainable urban management and environmental governance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 31

SDGBiasBench: Benchmarking and Mitigating Vision--Language Models' Biases in Sustainable Development Goals

SDGBiasBench is a large-scale benchmark suite designed to evaluate and mitigate biases in vision–language models (VLMs) when reasoning about Sustainable Development Goals (SDGs). It contains 500k expert‑involved multiple‑choice questions and 50k regression tasks, allowing assessment of both decision‑level and estimation‑level bias. Experiments show that current VLMs exhibit intrinsic SDG bias, often relying on priors rather than multimodal evidence, and the proposed CADE method significantly reduces this bias, improving accuracy and reducing mean absolute error.

By Zihang Lin, Huaiyuan Qin, Muli Yang, Hongyuan Zhu
arXiv Computer Vision
Aug 25

WADE: A Reasoning-Annotated Benchmark for Multi-Instance Floating-Waste Grounding with Compact Vision-Language Models

arXiv:2608.22950v1 Announce Type: new Abstract: Floating waste in inland waterways threatens aquatic ecosystems and requires timely monitoring under cluttered, multi-object conditions. Existing aquat...

By Md. Asaduzzaman Shuvo, Ahsan Farabi, Md. Abdul Ahad Minhaz, Mahedi Hasan, Israt Khandaker, Ibrahim Khalil Shanto, Muhammad Nomani Kabir
arXiv AI
Sep 18

A Multi-Modal Generative Model for Tomato Disease Leaves Understanding

The paper introduces SOLAR, a multimodal generative model that jointly interprets visual and textual data to understand tomato leaf diseases across six question‑answering tasks. SOLAR aligns visual features with task‑aware language representations using a Fusion Expert module based on a mixture‑of‑experts, enabling it to generate contextually relevant answers for diverse diagnostic tasks. Evaluated on 41,677 images and 216,209 QA pairs, SOLAR outperforms state‑of‑the‑art vision‑only, vision‑language, and task‑specific models in both closed and open‑ended settings, demonstrating superior accuracy, robustness, and multimodal reasoning.

By Khang Nguyen Quoc, Minh-Phuoc Tran, Gia-Han Truong, Luyl-Da Quach