Research data repositories are essential infrastructure for scientific inquiry and for ensuring that datasets follow FAIR (Findable, Accessible, Interoperable, and Reusable) principles. However, repos...
The paper presents a knowledge graph built from Harvard Dataverse’s public data, linking 102,650 datasets to 215,985 nodes and 528,003 edges that include keywords, publications, subjects, journals, and locations. About 43,991 datasets contain geospatial metadata, and 7,654 are identified as policy‑relevant, with elections and legislatures forming the largest cluster. The authors highlight the challenge of place resolution—disconnected nodes representing the same location—and propose the graph as a testbed for AI‑driven metadata enrichment and entity resolution, noting a bias toward American city‑level data.
By Danny EBanks, Devika Jain
The paper introduces an end‑to‑end pipeline that harvests, validates, and models links between scholarly articles and their source code from journals such as JOSS, SoftwareX, and IPOL, as well as SIGMOD ARI reproducibility reports. It produces a curated set of 4,397 DOI‑repository pairs and defines two Wikidata‑based application profiles—one for articles and one for software—aligned with schema.org and CodeMeta. Using these profiles, the authors created 4,182 new Wikidata software items linked to their papers, while only 82 repositories were previously represented, and they demonstrate compatibility with the COAR Notify protocol for future live enrichment.
By Camillo Carlo Pellizzari di San Girolamo, Francesco Tosoni
arXiv:2607. 05970v1 Announce Type: cross Abstract: Dataset search depends heavily on metadata, making LLM-generated metadata a consequential form of synthetic content in retrieval systems.
By Riccardo Terrenzi, Serkan Ayvaz
Scientific research increasingly relies on large, heterogeneous data sources, motivating interest in retrieval-augmented generation (RAG) systems that provide natural language access to scientific kno...
arXiv:2606. 07611v1 Announce Type: cross Abstract: This paper proposes an improved approach to the analysis of Mining Software Repositories (MSR) datasets via metadata enrichment, FAIRness assessment, and topic-driven analysis.
By Aabia Ather, Muhammad Usayd Ather, Qurat-Ul-Ain Somroo, Muhammad Khuram Shahzad
arXiv:2609.16519v1 Announce Type: new
Abstract: Scientific research increasingly relies on large, heterogeneous data sources, motivating interest in retrieval-augmented generation (RAG) systems that...
By Bernie Boscoe, Srinath Saikrishnan, Vikram Seenivasan, Jack Stark, Andrew Lizarraga, Morgan Himes, Jonathan Soriano, PJ Allen, Tuan Do
arXiv:2608. 19625v1 Announce Type: new Abstract: Scientific data are increasingly used by AI agents, yet existing dataset representations provide limited support for autonomous discovery, interpretation, and invocation.
By Xiaohan Huang, Qingqing Long, Xiaolei Du, Siyu Pu, Jiawen Xu, Haotian Chen, Chenyang Zhao, Jinbiao Liu, Xuezhi Wang, Hao Wang, Hengshu Zhu, Yuanchun Zhou
The FAIR Digital Object (FDO) framework mandates that metadata attribute values be expressed as persistent identifiers (PIDs) wherever possible, to produce a fully machine-actionable graph in which ev...
arXiv:2608.23263v1 Announce Type: new
Abstract: The FAIR Digital Object (FDO) framework mandates that metadata attribute values be expressed as persistent identifiers (PIDs) wherever possible, to pro...
By Zeyd Boukhers, Lingxiao Kong, Xenophon Zabulis, Georgios Toubekis
The paper introduces the Scientific Data Skill (SciDSK), an agent‑ready representation that packages dataset‑specific knowledge and operational guidance as a reusable skill. SciDSK integrates dataset descriptions, scientific context, file organization, usage procedures, quality checks, and provenance information while keeping the data in its original repository. The authors define a structured specification, build a construction pipeline, and launch the Scientific Data Skill Bank to publish SciDSK resources across six scientific disciplines, demonstrating improved agent‑driven dataset discovery and interpretation through evaluation benchmarks.
arXiv:2608. 07254v1 Announce Type: cross Abstract: The increasing specialization of scientific research challenges existing classification systems, which provide effective representations of broad disciplines and research topics but often fail to capture the fine-grained conceptual structure of contemporary science.
By Daniele Raimondi, Feichi Lu, Oliver Grun, Mariia Eremina, Andrea Perlato