The paper introduces a pipeline that merges structured disaster records from EM‑DAT with unstructured documents from ReliefWeb and the European Media Monitor to generate source‑grounded disaster storylines and causal knowledge graphs. Using Retrieval‑Augmented Generation, it produces tabular event profiles covering 17 fields and builds causal graphs enriched with citation‑grounded explanatory narratives, allowing traceability to primary sources. Human evaluation across three crisis cases shows high retrieval precision, strong faithfulness of causal relations, and a clear expert preference for citation‑grounded components over ungrounded ones.
By Ivan Decostanzi, Michele Ronco, Sergio Consoli, Christina Corbane, Lorenzo Bertolini, Indaco Biazzo, Daria Mihaila, Manuel Garcia-Herranz, Felix Schwebel, Yelena Mejova, Kyriaki Kalimeri
arXiv:2607. 03447v1 Announce Type: cross Abstract: Knowledge graphs (KGs) that underpin Graph-based Retrieval-Augmented Generation (Graph-RAG) are increasingly built automatically by LLM-driven extraction rather than curated by experts.
By Axel TahmasebiMoradi, Lucas Schott, Martin Royer
arXiv:2607. 02387v1 Announce Type: cross Abstract: NASA and its data centers hold thousands of geoscience datasets and tools like Worldview, Giovanni, the Science Discovery Engine, and Harmony.
By Minghan Yu, Youran Sun, Chugang Yi, Yixin Wen, Haizhao Yang
arXiv:2511.04473v3 Announce Type: replace
Abstract: Retrieval of information from graph-structured knowledge bases represents a promising direction for improving the factuality of LLMs. While various...
By Alberto Cattaneo, Carlo Luschi, Daniel Justus
HyGRAIL is a framework for discovering scientific hypotheses in incomplete knowledge graphs by combining a graph neural network (GNN) triage with large language model (LLM) review. The GNN scores candidate hypotheses and routes only ambiguous cases to the LLM, which receives structured evidence from the graph converted into natural language. Experiments on MatKG show HyGRAIL achieves the highest F1 score, improves over baselines, and cuts LLM calls by over 54%.
By Yihang Sun, Zhihan Zhu, Zhiyuan Jiang, Jingyi Ge, Zixuan Li, Jiaxuan You
arXiv:2603.28773v2 Announce Type: replace-cross
Abstract: Large language models (LLMs) frequently generate confident yet factually incorrect content when used for language generation (a phenomenon of...
By Dobrik Georgiev, Kheeran K. Naidu, Alberto Cattaneo, Federico Monti, Carlo Luschi, Daniel Justus
The paper introduces EXYGEN, a framework that enables conversational access to large knowledge graphs by combining VoID descriptions, ShEx schemas, retrieved triples, and example question‑query pairs in a retrieval‑augmented generation pipeline. On the SciQA benchmark, this approach achieves an exact‑match score of 0.419 without fine‑tuning any large language model, and shows that larger general‑purpose LLMs can outperform smaller code‑specialized ones when provided sufficient context. To scale metadata generation for very large KGs, the authors propose a predicate‑coverage‑aware parallel graph sampling strategy that preserves structural diversity, reduces runtime by over 80× on OpenCitations Meta and GESIS, and is the only tractable method for obtaining complete metadata on ORKG.
By Harshdeep Singh, Yurui Zhu, Giovanni Colavizza, Matteo Romanello
arXiv:2603.21152v4 Announce Type: replace-cross
Abstract: Modern seismic networks resolve earthquake sequences in unprecedented detail, yet explaining how large earthquakes emerge from evolving fault...
By Feng Liu, Xin Cui, Jian Xu, Xinghao Wang, Zijie Guo, Jiong Wang, S. Mostafa Mousavi, Xinyu Gu, Hao Chen, Ben Fei, Lihua Fang, Fenghua Ling, Zefeng Li, Lei Bai
arXiv:2608. 09276v1 Announce Type: cross Abstract: Planetary geology relies on historical, interpretive reasoning to reconstruct past events from diverse observations.
By Tom Sander, Kay Wohlfarth, Christian W\"ohler
The paper introduces AHLERT, a system that automatically extracts environment-aware hunt leads from Cyber Threat Intelligence reports. It combines a hybrid retriever—dense vector search plus multi-hop knowledge‑graph traversal seeded with MITRE ATT&CK—with ontology‑grounded retrieval‑augmented generation to constrain leads to a defender’s assets. Evaluations on public CTI reports show that AHLERT doubles mean F1 scores and achieves an effectiveness score of ~86.95% compared to off‑the‑shelf LLM models.
By Akash Prakash, Boubakr Nour, Makan Pourzandi, Chadi Assi, Mourad Debbabi
Build2SPARQL is a large-scale benchmark dataset for translating natural-language questions into SPARQL queries over building knowledge graphs. The dataset is generated by a KG‑grounded pipeline that produces 6,136 executable SPARQL queries and 30,680 corresponding natural-language questions across six query-pattern families and five vocabulary registers, covering 201 building KGs. Human validation shows high semantic fidelity, naturalness, and operational plausibility, and retrieval‑augmented evaluation demonstrates significant accuracy gains for open‑weight language models.
By Wooyoung Jung
arXiv:2606. 05415v1 Announce Type: cross Abstract: Real-world data spans tables, documents, and semi-structured files with implicit semantics.
By Padmaja Jonnalagedda, Yuguang Yao, Xiang Gao, Hilaf Hasson, Kamalika Das