arXiv AI By Zhangding Liu, Neda Mohammadi, John E. Taylor

Knowledge-Guided Vision-Language Inference for Image-Based Urban Flood Depth Estimation

Read the original on arXiv AI →

The paper introduces FloodVision, a knowledge-guided vision‑language framework that estimates urban flood depth from a single RGB image. It combines a general‑purpose vision‑language model with FloodKG, a domain knowledge base that encodes canonical object dimensions and component landmarks to promote component‑level reasoning. On 654 crowdsourced New York flood images, FloodVision cuts mean absolute error from 15.62 cm to 8.75 cm and median error from 14.35 cm to 7.75 cm, outperforming the VLM‑only baseline in 69.3 % of cases.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jun 4

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models

Multimodal Large Language Models (MLLMs) excel at 2D semantic understanding but lack intrinsic 3D awareness, resulting in representations that fail to maintain geometric and spatial consistency across video frames. Given the scarcity of large-scale 3D data, we present GeoVR, a novel framework that learns geometric representations using purely 2D video sequences.

arXiv AI
Aug 18

FloodReasonBench: Benchmarking VLM Reasoning Segmentation for Embodied Flood Response at the Edge

arXiv:2608. 15410v1 Announce Type: cross Abstract: Reasoning segmentation enables vision-language models (VLMs) to translate mission-relevant language requests into pixel-level visual grounding, offering a natural perception interface for embodied agents.

By Rajat Bhattacharjya, Yoomee Jung, Minwoo Kim, Sing-Yao Wu, Eli Bozorgzadeh, Nalini Venkatasubramanian, Nikil Dutt