The paper introduces an end‑to‑end pipeline that harvests, validates, and models links between scholarly articles and their source code from journals such as JOSS, SoftwareX, and IPOL, as well as SIGMOD ARI reproducibility reports. It produces a curated set of 4,397 DOI‑repository pairs and defines two Wikidata‑based application profiles—one for articles and one for software—aligned with schema.org and CodeMeta. Using these profiles, the authors created 4,182 new Wikidata software items linked to their papers, while only 82 repositories were previously represented, and they demonstrate compatibility with the COAR Notify protocol for future live enrichment.
By Camillo Carlo Pellizzari di San Girolamo, Francesco Tosoni
arXiv:2609.16498v1 Announce Type: cross
Abstract: Research data repositories are essential infrastructure for scientific inquiry and for ensuring that datasets follow FAIR (Findable, Accessible, Inte...
By Daniel Ebanks, Devika Jain
Research data repositories are essential infrastructure for scientific inquiry and for ensuring that datasets follow FAIR (Findable, Accessible, Interoperable, and Reusable) principles. However, repos...
The paper introduces CodeGraph, an open‑taxonomy knowledge graph that semantically annotates source code by extracting entities such as algorithms, paradigms, design patterns, and application domains from millions of files. Using a specialized large language model and a three‑stage Wikidata linking process, the authors ground these entities in Wikidata and construct a graph with about 158 million nodes and 1 billion typed edges across 14 programming languages. A quality‑assurance protocol combining human evaluation and an LLM‑as‑a‑judge filter quantifies annotation precision.
By Federico Pennino, Andrea Gurioli, Stefano Zacchiroli, Maurizio Gabbrielli, Paolo Ferragina
The paper reports on building a research-software catalog using a coding agent, starting from a three‑day hackathon prototype and moving to public deployment. It details the engineering work needed—adversarial review, data‑quality checks, browser validation, and publication safeguards—to ensure reliable operation, noting that silent failures were more problematic than crashes. The authors then examine applying these lessons to a larger, human‑curated portal (MateriApps) that combines curated metadata, external documentation, vector search, and local language‑model generation, finding that explicit validation, monitoring, and repeated review remain essential for AI‑assisted software portals.
By Kazuyoshi Yoshimi, Satoshi Terasaki, Gotai Yamada
arXiv:2606. 30304v1 Announce Type: cross Abstract: This paper presents preliminary findings from a UKRI-funded Metascience project comparing three LLM-based approaches, GPT-4o, Mistral, and a bespoke algorithm, DSIT-Taxonomies, for extracting and classifying research entities from funding proposals.
By Xingran Ruan, Angelo Salatino, Rosa Filgueira, Kara Moraw, Alexandru Marcoci, Gemma Derrick, Sarah Callaghan
arXiv:2606. 24894v2 Announce Type: replace-cross Abstract: Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited.
By Anzhe Xie, Weihang Su, Jiaxin Mao, Yiqun Liu, Shaoping Ma, Qingyao Ai
arXiv:2608. 19625v1 Announce Type: new Abstract: Scientific data are increasingly used by AI agents, yet existing dataset representations provide limited support for autonomous discovery, interpretation, and invocation.
By Xiaohan Huang, Qingqing Long, Xiaolei Du, Siyu Pu, Jiawen Xu, Haotian Chen, Chenyang Zhao, Jinbiao Liu, Xuezhi Wang, Hao Wang, Hengshu Zhu, Yuanchun Zhou
arXiv:2609.07586v1 Announce Type: new
Abstract: Software repositories contain vast amounts of data on code contributions, bug reports, and project activities, yet this information remains challenging...
By Muhammad Jawad Chowdhury, Md. Sakib Khan
W-RAG is a source-aware retrieval framework designed for enterprise document generation from heterogeneous knowledge bases. It uses ontology-guided retrieval, local ranking within each knowledge base, and source-level weighting to balance evidence from diverse sources. A new dataset covering multiple document types and industry domains demonstrates that W-RAG improves document coverage and generation quality compared to standard RAG pipelines.
By Hridya Dhulipala, Rajesh Ombase, Michael Wang, Tien N. Nguyen
The paper introduces the Scientific Data Skill (SciDSK), an agent‑ready representation that packages dataset‑specific knowledge and operational guidance as a reusable skill. SciDSK integrates dataset descriptions, scientific context, file organization, usage procedures, quality checks, and provenance information while keeping the data in its original repository. The authors define a structured specification, build a construction pipeline, and launch the Scientific Data Skill Bank to publish SciDSK resources across six scientific disciplines, demonstrating improved agent‑driven dataset discovery and interpretation through evaluation benchmarks.
arXiv:2605.29522v2 Announce Type: replace
Abstract: As scientific literature grows rapidly and research increasingly involves AI agents, automated survey generation has become a key capability for bo...
By Ziyue Yang, Da Ma, Hanqi Li, Zijian Wang, Tiancheng Huang, Zijian Hu, Chenrun Wang, Yunzhe Zhang, Xiaobao Wu, Kai Yu, Lu Chen