arXiv:2608. 00012v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated.
By Fengxiang Wang, Qiuyang Yu, Yueying Li, Mingshuo Chen, Chengchi Fei, Kaiyi Xu, Lixin Gu, Wangxu Wei, Junchao Gong, Lipeng Ma, Jiong Wang, Fenghua Ling, Wenlong Zhang, Xue Yang, Wenjing Yang, Ben Fei, Long Lan
arXiv:2607. 24588v1 Announce Type: new Abstract: Early warning of extreme weather is essential for mitigating the societal, economic, and environmental risks posed by hazardous weather events.
By Hang Ni, Weijia Zhang, Fan Liu, Mengqian Lu, Hao Liu
arXiv:2607. 16050v1 Announce Type: new Abstract: Pluvial (rainfall-driven) flooding accounts for 45% of National Flood Insurance Program (NFIP) claims in the United States and is harder to predict than its riverine and coastal counterparts, with existing approaches limited to coarse resolution, regional domains, or computationally intensive process-based models unsuitable for daily continental-scale use.
By Yuya Kawakami, Daniel Cayan, Dongyu Liu, Kwan-Liu Ma, Tom Corringham
The paper introduces SmartWeatherAgent, a three‑stage architecture that combines intent recognition, hazard prediction, and reasoning‑enhanced generation to improve tourism meteorological services. It fuses rule‑based methods with large language models and a LightGBM model enriched with highland‑specific features, achieving an F1‑Macro score of 0.605 and 1.60 ms latency on high‑wind, precipitation, and low‑temperature events. A 12‑round micro‑step prompt self‑optimization loop raises the composite warning quality score from 4.2 to 8.9, with notable gains in data source citation, physical mechanism explanation, and scientific rigor through explicit uncertainty statements.
By Shuai Yan, Yang Xu, Shan He
SeisEvo is a method that uses a large language model (LLM) and multi‑agent search to evolve seismic data reconstruction algorithms rather than optimize a single result. Starting from a classical algorithm, the agents modify only user‑opened components, rejecting candidates that violate physical constraints and scoring the rest by execution. The resulting white‑box algorithms—such as a residual‑gated, phase‑aligned dip‑consistency projection for interpolation and a reliability‑grouped singular‑value shrinkage for simultaneous interpolation and denoising—outperform classic methods by several decibels and generalize to unseen data.
By Yingjie Xu, Siwei Yu, Jianwei Ma
To address insufficient contextualization, weak generalization, and poor scenario adaptation in tourism meteorological services, we propose SmartWeatherAgent--a unified three-stage architecture integr...
Geographic Information System (GIS) professionals rely on multi-step spatial analysis workflows to support decision-making in urban planning, disaster response, and environmental monitoring. The process is tedious, time-consuming, and error-prone.
PermitGPT is a generative‑AI framework that transforms unstructured construction permit descriptions into structured outputs for safety hazard identification, permit requirement specification, and community impact assessment. It aligns data from the NYC Department of Buildings, OSHA, and NYC 311 to create 90,000 prompt‑response pairs, fine‑tunes three open‑weight language models, and evaluates them on 2,833 test cases, reporting complementary performance across inference speed, lexical overlap, and semantic alignment. The study presents an initial AI‑assisted approach to construction governance and outlines future evaluation and validation directions.
By Mohd Ruhul Ameen, Farjana Aktar, Akif Islam, Momen Khandoker Ope, Abu Saleh Musa Miah, Jungpil Shin
arXiv:2604. 02022v4 Announce Type: replace Abstract: Evaluating the safety of LLM-based agents is increasingly important because risks in realistic deployments often emerge over multi-step interactions rather than isolated prompts or final responses.
By Yu Li, Haoyu Luo, Yuejin Xie, Yuqian Fu, Zhonghao Yang, Shuai Shao, Qihan Ren, Wanying Qu, Yanwei Fu, Yujiu Yang, Jing Shao, Xia Hu, Dongrui Liu
arXiv:2604. 12306v3 Announce Type: replace-cross Abstract: Climate decision-making in the GCC states increasingly demands systems that can translate heterogeneous scientific and policy evidence into actionable guidance, yet general-purpose large language models (LLMs) remain weak both in region-specific climate knowledge and grounded interaction with geospatial and forecasting tools.
By Muhammad Umer Sheikh, Khawar Shehzad, Salman Khan, Fahad Shahbaz Khan, Muhammad Haris Khan
SAFARI is the first industrial benchmark for evaluating large language models (LLMs) in automotive hazard analysis and risk assessment (HARA) under ISO 26262. It comprises 3,000 de‑identified HARA cases and tests two tasks: open‑ended hazard generation and standards‑grounded risk classification, using a novel reference‑anchored LLM‑as‑a‑judge protocol. Experiments with nine state‑of‑the‑art LLMs show that while hazard narratives are often plausible, risk classification remains weak (best ASIL macro‑F1 = 0.261), with errors mainly due to missing scenario context and misjudged controllability.
"whyItMatters":"The benchmark highlights the current limitations of LLMs in safety‑critical engineering workflows, guiding future research and expert oversight in automotive safety analysis."
By Chenxi Wu, Zimu Wang, Haiyang Zhang, Wei Wang, Zhijie Xu
arXiv:2607. 17437v1 Announce Type: new Abstract: Large language model (LLM) agents offer a generative approach to simulating human behavior under conditions that may have few or no direct historical analogues, a common challenge in disaster and infrastructure-disruption planning.
By Chen Xia, Zexi Kuang, Yuqing Hu