arXiv:2607.12441v3 Announce Type: replace
Abstract: Wikipedia plays a key role in shaping public understanding of science, and its openly accessible revision history is a unique record of how scienti...
By Omer Ehrlich, Nitzan Barzilay, Rona Aviram, Tom Hope
The paper introduces ingest‑time fact compilation, an architecture that preprocesses and compiles corpus data into self‑contained facts with resolved revisions, deletions, and source trust. By storing this compiled state, query‑time models can retrieve answers directly, avoiding costly reconstruction from raw passages. Experiments show that this approach reduces read cost per question by 12.89× and token usage by 21.6× while maintaining accuracy.
By Kyle Wild, Yusuke Takahashi, Asako Uraki
The paper introduces an end‑to‑end pipeline that harvests, validates, and models links between scholarly articles and their source code from journals such as JOSS, SoftwareX, and IPOL, as well as SIGMOD ARI reproducibility reports. It produces a curated set of 4,397 DOI‑repository pairs and defines two Wikidata‑based application profiles—one for articles and one for software—aligned with schema.org and CodeMeta. Using these profiles, the authors created 4,182 new Wikidata software items linked to their papers, while only 82 repositories were previously represented, and they demonstrate compatibility with the COAR Notify protocol for future live enrichment.
By Camillo Carlo Pellizzari di San Girolamo, Francesco Tosoni
arXiv:2609.37226v1 Announce Type: cross
Abstract: Answering questions and completing tasks over large document collections often requires connecting evidence spread across multiple documents, such as...
By Soyeong Jeong, Sujay Kumar Jauhar, Sung Ju Hwang, Andrew Joohun Nam
Scientific abstracts mix contributions with background, motivation, and meta-language, so tools that read them as-is cannot separate what a field produces from what it discusses. We present Drift Insp...
Drift Inspector is an open‑source system that extracts Atomic Contribution Claims (ACCs) from scientific abstracts using an LLM, then clusters these claims over time to map how a research field evolves. Applied to six years of EMNLP, the tool reveals a shift from classic NLP tasks toward LLM‑era capabilities such as reasoning and multimodality—trends that keyword or whole‑abstract counts miss. The pipeline has also processed the entire ACL Anthology, yielding 346,000 claims from 80,000 abstracts across 423 venues, with human‑validated extraction and clustering aligned to an external taxonomy.
By Vsevolod Karimov, Stepan Ostarkov, Anastasia Poroshina, Anatoly Frolov, Alexander Panchenko