arXiv:2609.16498v1 Announce Type: cross
Abstract: Research data repositories are essential infrastructure for scientific inquiry and for ensuring that datasets follow FAIR (Findable, Accessible, Inte...
By Daniel Ebanks, Devika Jain
The paper presents a knowledge graph built from Harvard Dataverse’s public data, linking 102,650 datasets to 215,985 nodes and 528,003 edges that include keywords, publications, subjects, journals, and locations. About 43,991 datasets contain geospatial metadata, and 7,654 are identified as policy‑relevant, with elections and legislatures forming the largest cluster. The authors highlight the challenge of place resolution—disconnected nodes representing the same location—and propose the graph as a testbed for AI‑driven metadata enrichment and entity resolution, noting a bias toward American city‑level data.
By Danny EBanks, Devika Jain
Scientific research increasingly relies on large, heterogeneous data sources, motivating interest in retrieval-augmented generation (RAG) systems that provide natural language access to scientific kno...
arXiv:2607. 05970v1 Announce Type: cross Abstract: Dataset search depends heavily on metadata, making LLM-generated metadata a consequential form of synthetic content in retrieval systems.
By Riccardo Terrenzi, Serkan Ayvaz
The paper introduces an end‑to‑end pipeline that harvests, validates, and models links between scholarly articles and their source code from journals such as JOSS, SoftwareX, and IPOL, as well as SIGMOD ARI reproducibility reports. It produces a curated set of 4,397 DOI‑repository pairs and defines two Wikidata‑based application profiles—one for articles and one for software—aligned with schema.org and CodeMeta. Using these profiles, the authors created 4,182 new Wikidata software items linked to their papers, while only 82 repositories were previously represented, and they demonstrate compatibility with the COAR Notify protocol for future live enrichment.
By Camillo Carlo Pellizzari di San Girolamo, Francesco Tosoni
arXiv:2609.16519v1 Announce Type: new
Abstract: Scientific research increasingly relies on large, heterogeneous data sources, motivating interest in retrieval-augmented generation (RAG) systems that...
By Bernie Boscoe, Srinath Saikrishnan, Vikram Seenivasan, Jack Stark, Andrew Lizarraga, Morgan Himes, Jonathan Soriano, PJ Allen, Tuan Do
The FAIR Digital Object (FDO) framework mandates that metadata attribute values be expressed as persistent identifiers (PIDs) wherever possible, to produce a fully machine-actionable graph in which ev...
arXiv:2608.23263v1 Announce Type: new
Abstract: The FAIR Digital Object (FDO) framework mandates that metadata attribute values be expressed as persistent identifiers (PIDs) wherever possible, to pro...
By Zeyd Boukhers, Lingxiao Kong, Xenophon Zabulis, Georgios Toubekis
The zbMATH Open Knowledge Graph is a large-scale RDF knowledge graph that spans more than 250 years of mathematical scholarship. It goes beyond traditional bibliographic metadata by incorporating expert-curated semantic content such as reviews, keywords, subject classifications, software references, and disambiguated authorship. With 34 million entities and 168 million RDF triples, the graph enables fine-grained, historically grounded exploration of mathematical concepts, research fields, and scholarly relationships over time.
By Yuni Susanti, Moritz Schubotz
arXiv:2608. 07254v1 Announce Type: cross Abstract: The increasing specialization of scientific research challenges existing classification systems, which provide effective representations of broad disciplines and research topics but often fail to capture the fine-grained conceptual structure of contemporary science.
By Daniele Raimondi, Feichi Lu, Oliver Grun, Mariia Eremina, Andrea Perlato
arXiv:2608. 19625v1 Announce Type: new Abstract: Scientific data are increasingly used by AI agents, yet existing dataset representations provide limited support for autonomous discovery, interpretation, and invocation.
By Xiaohan Huang, Qingqing Long, Xiaolei Du, Siyu Pu, Jiawen Xu, Haotian Chen, Chenyang Zhao, Jinbiao Liu, Xuezhi Wang, Hao Wang, Hengshu Zhu, Yuanchun Zhou
arXiv:2606. 07611v1 Announce Type: cross Abstract: This paper proposes an improved approach to the analysis of Mining Software Repositories (MSR) datasets via metadata enrichment, FAIRness assessment, and topic-driven analysis.
By Aabia Ather, Muhammad Usayd Ather, Qurat-Ul-Ain Somroo, Muhammad Khuram Shahzad