The paper introduces a pipeline that merges structured disaster records from EM‑DAT with unstructured documents from ReliefWeb and the European Media Monitor to generate source‑grounded disaster storylines and causal knowledge graphs. Using Retrieval‑Augmented Generation, it produces tabular event profiles covering 17 fields and builds causal graphs enriched with citation‑grounded explanatory narratives, allowing traceability to primary sources. Human evaluation across three crisis cases shows high retrieval precision, strong faithfulness of causal relations, and a clear expert preference for citation‑grounded components over ungrounded ones.
By Ivan Decostanzi, Michele Ronco, Sergio Consoli, Christina Corbane, Lorenzo Bertolini, Indaco Biazzo, Daria Mihaila, Manuel Garcia-Herranz, Felix Schwebel, Yelena Mejova, Kyriaki Kalimeri
arXiv:2607. 22592v1 Announce Type: new Abstract: Graph-based retrieval-augmented generation (GraphRAG) grounds answers in structured knowledge, but current systems extract entities and relationships exhaustively, producing graphs whose size and construction cost scale with corpus length rather than with the reasoning a query requires.
By Marc Saouda (Boston Consulting Group), Rajprakash Bale (Boston Consulting Group), Eren Aldis (Boston Consulting Group), Cloves Almeida (Boston Consulting Group)
ConstructCIE is a manually annotated dataset designed for extracting causal information from OSHA construction accident reports. It employs a hierarchical schema that categorizes accident types, causal factors, sub‑causal factors, and the supporting evidence spans. Experiments with supervised sequence taggers and instruction‑tuned large language models show strong performance on accident‑type prediction and broad causal recovery, yet precise span‑level extraction remains challenging, highlighting the need for better domain grounding and evidence extraction.
By Hung Nguyen, Jaehoon Lee, Namgyun Kim, Kuan-Hao Huang
Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved promising performance in part of these tasks individually, their capability in event analysis still lacks comprehensive understanding due to restricted document granularity, task designs, and data source of existing benchmarks.
arXiv:2606. 07525v1 Announce Type: cross Abstract: Causal graphs in text are typically populated by observable, predefined events.
By Liesbeth Allein, Marie-Francine Moens
arXiv:2606. 19710v1 Announce Type: cross Abstract: Court proceedings contain valuable evidence about human smuggling networks, but this information is often buried within unstructured, jargon-heavy legal documents.
By Elijah Feldman, Dipak Meher, Carlotta Domeniconi
arXiv:2607. 09094v1 Announce Type: cross Abstract: Legal precedent retrieval is a fundamental task in legal case preparation, planning, litigation strategy, and legal research.
By Devanshu Verma, Vasudha Bhatnagar, Vikas Kumar, Balaji Ganesan
LexIssue introduces a benchmark for identifying disputed legal issues in Chinese civil litigation, comprising 430 real‑world cases and 1,303 expert‑annotated issues. The dataset is built around a hierarchical schema that links free‑form issue descriptions to structured legal categories, enabling two complementary tasks: issue generation and issue classification. A retrieval‑augmented knowledge base covering 27 causes of action and 441 issue entries is provided, and experiments show that incorporating this knowledge consistently improves model performance on the tasks.
By Huiyuan Xie, Yuqin Huang, Zhicheng Hao, Yida Cai, Shaochun Wang, Zhenghao Liu, Yuxiao Ye
arXiv:2511. 03217v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel in generating fluent utterances but can lack reliable grounding in verified information.
By Shaghayegh Kolli, Richard Rosenbaum, Timo Cavelius, Lasse Strothe, Andrii Lata, Jana Diesner
Humanitarian reports are long, noisy, and multi-topic, making it difficult to consolidate decision-relevant causal evidence. We present a ReliefWeb study (2000-2024) and a two-stage Large Language Model (LLM) pipeline that extracts structured intervention-outcome records with direction and strength attributes.
EvidenceNet is a disease‑specific dataset that transforms full‑text biomedical literature into structured evidence records and graph representations, preserving study design, provenance, and quantitative support. Using an LLM‑assisted pipeline, it extracts experimentally grounded findings, normalizes entities, scores evidence quality, and links related records via typed semantic relations. The released subsets—EvidenceNet‑HCC and EvidenceNet‑CRC—contain thousands of evidence records and richly connected graphs, with high extraction and relation‑type accuracy, enabling retrieval‑augmented question answering and graph‑based tasks such as link prediction and target prioritization.
By Chang Zong, Jinyu Chen, Sicheng Lv, Si-tu Xue, Huilin Zheng, Jian Wan, Lei Zhang
Legal argument mining supports passage classification, retrieval, and argument completion. This work introduces an expert-annotated dataset of 42 U.S. federal tax opinions on corporate reorganizations...