arXiv AI By Zhuoran Li, Weiyi Kong, Boer Zhang

CDEP Agent: Connecting Meteorologically Detected Temporal Compound Events to Real-World Documentary Evidence

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv AI
Jun 3

Auditable Climate Risk Intelligence from Fragmented ESG Data: Deterministic Orchestration and Imbalance-Aware Learning for Scope 1-3 Validation

arXiv:2606. 02604v1 Announce Type: cross Abstract: ESG and climate risk data remain fragmented across heterogeneous Scope 1, Scope 2, and Scope 3 reporting environments, while conventional validation pipelines lack provenance aware auditability, hidden drift detection, and reproducibility oriented governance.

By Karan Sehgal, Khawar Naveed Bhatti
arXiv Machine Learning
Aug 27

How Robust Are Automated Fact-Checking Systems? A Cross-Benchmark Evaluation

The paper evaluates the robustness of automated fact‑checking systems by cross‑benchmarking nine models—including random baselines, fine‑tuned transformers, zero‑shot LLMs, and top AVeriTeC 2025 systems—across four datasets from scientific, open‑web, and climate domains. It finds that fine‑tuned models outperform zero‑shot LLMs on ClimateCheck, that system rankings vary strongly with domain and metric, and that replacing retrieved evidence with gold annotations boosts veracity accuracy by 14–22 points, underscoring retrieval as the main bottleneck. The authors provide code, pre‑processed datasets, and results to enable reproducible research.

By Aida Usmanova, Zangir Iklassov, Markus Leippold, Ricardo Usbeck
arXiv Computation and Language
6d ago

Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations

The paper introduces ARGUS, a language‑model pipeline that audits evidence for identification assumptions in difference‑in‑differences studies of climate policy. ARGUS evaluates reported evidence against an eleven‑dimension rubric, abstaining when evidence cannot be retrieved. In tests, ARGUS detects 73% of injected flaws versus 18% for a keyword approach, abstains on about 40% of assessments in 26 economics papers, and often assigns higher risk than human labels in a five‑paper pilot.

By Yonghong Zhang, Yong Xie, Isabel M. Parra, Ricardo Correia
arXiv AI
Aug 25

EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards

arXiv:2608.23525v1 Announce Type: new Abstract: Earth-system analysis reconstructs changing physical processes from observations that differ in source, scale, timing, and modality. Natural hazards ma...

By Zhiqing Cui, Xinxiang Yin, Yihong Tang, Xinglang Zhang, Yuanzhe Hu, Siru Zhong, Weidong Tang, Yuxuan Liang, Weijia Li, Ming Jin, Shirui Pan, Yuhao Kang, Dingyi Zhuang, Jinhua Zhao
arXiv Machine Learning
Aug 14

Which Site, and When: A Free-Satellite-Data Test of Himalayan Glacial Lake Bursts, Landslides, and Ice Floods

arXiv:2608. 12422v1 Announce Type: new Abstract: Two free satellite signals carry real information about glacial-lake outburst risk in the Nepal Himalaya: radar interferometry sees a moraine dam slowly sagging, and satellite weather marks the weeks when a primed lake is under stress.

By Matthew Kahn, Milan Arjel, Nirmala Adhikari, Mingmar Sherpa, James Pope