arXiv AI By Aabia Ather, Muhammad Usayd Ather, Qurat-Ul-Ain Somroo, Muhammad Khuram Shahzad

MIRAGE: Metadata-Integrated Repository Analysis and Guided Enhancement for MSR Datasets

Read the original on arXiv AI →

arXiv:2606. 07611v1 Announce Type: cross Abstract: This paper proposes an improved approach to the analysis of Mining Software Repositories (MSR) datasets via metadata enrichment, FAIRness assessment, and topic-driven analysis.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 21

From Code Archival to Knowledge Graph: Bridging Software Heritage, COAR Notify and Wikidata

The paper introduces an end‑to‑end pipeline that harvests, validates, and models links between scholarly articles and their source code from journals such as JOSS, SoftwareX, and IPOL, as well as SIGMOD ARI reproducibility reports. It produces a curated set of 4,397 DOI‑repository pairs and defines two Wikidata‑based application profiles—one for articles and one for software—aligned with schema.org and CodeMeta. Using these profiles, the authors created 4,182 new Wikidata software items linked to their papers, while only 82 repositories were previously represented, and they demonstrate compatibility with the COAR Notify protocol for future live enrichment.

By Camillo Carlo Pellizzari di San Girolamo, Francesco Tosoni
arXiv Computation and Language
Sep 25

CodeGraph: Open-Taxonomy Knowledge Graph for Source Code with Wikidata Grounding

The paper introduces CodeGraph, an open‑taxonomy knowledge graph that semantically annotates source code by extracting entities such as algorithms, paradigms, design patterns, and application domains from millions of files. Using a specialized large language model and a three‑stage Wikidata linking process, the authors ground these entities in Wikidata and construct a graph with about 158 million nodes and 1 billion typed edges across 14 programming languages. A quality‑assurance protocol combining human evaluation and an LLM‑as‑a‑judge filter quantifies annotation precision.

By Federico Pennino, Andrea Gurioli, Stefano Zacchiroli, Maurizio Gabbrielli, Paolo Ferragina
arXiv AI
Sep 7

Building a research-software catalog with a coding agent: from hackathon prototype to public deployment

The paper reports on building a research-software catalog using a coding agent, starting from a three‑day hackathon prototype and moving to public deployment. It details the engineering work needed—adversarial review, data‑quality checks, browser validation, and publication safeguards—to ensure reliable operation, noting that silent failures were more problematic than crashes. The authors then examine applying these lessons to a larger, human‑curated portal (MateriApps) that combines curated metadata, external documentation, vector search, and local language‑model generation, finding that explicit validation, monitoring, and repeated review remain essential for AI‑assisted software portals.

By Kazuyoshi Yoshimi, Satoshi Terasaki, Gotai Yamada
arXiv AI
Jun 30

Research Entity Extraction and Topic Detection from UKRI Grant Proposals

arXiv:2606. 30304v1 Announce Type: cross Abstract: This paper presents preliminary findings from a UKRI-funded Metascience project comparing three LLM-based approaches, GPT-4o, Mistral, and a bespoke algorithm, DSIT-Taxonomies, for extracting and classifying research entities from funding proposals.

By Xingran Ruan, Angelo Salatino, Rosa Filgueira, Kara Moraw, Alexandru Marcoci, Gemma Derrick, Sarah Callaghan