Drift Inspector is an open‑source system that extracts Atomic Contribution Claims (ACCs) from scientific abstracts using an LLM, then clusters these claims over time to map how a research field evolves. Applied to six years of EMNLP, the tool reveals a shift from classic NLP tasks toward LLM‑era capabilities such as reasoning and multimodality—trends that keyword or whole‑abstract counts miss. The pipeline has also processed the entire ACL Anthology, yielding 346,000 claims from 80,000 abstracts across 423 venues, with human‑validated extraction and clustering aligned to an external taxonomy.
By Vsevolod Karimov, Stepan Ostarkov, Anastasia Poroshina, Anatoly Frolov, Alexander Panchenko
How does research evolve, and what substrate would let us forecast where it goes next? Scientific progress is not simply a uniform accumulation of facts: ideas extend prior methods, address known limitations, realize proposed future directions, and sometimes dispute earlier claims.
arXiv:2606.22342v2 Announce Type: replace
Abstract: How does research evolve, and can we trace it at the level of individual claims? Scientific progress is not simply a uniform accumulation of facts....
By Abdul Muntakim, Md Abdullah Al Hafiz Khan, Sadid Hasan, Yong Pei
arXiv:2606. 23989v1 Announce Type: cross Abstract: End-to-end large language models (LLMs) produce fluent multi-document summaries but remain prone to hallucination, and the attributions they offer are typically coarse (whole documents or passages) and generated post hoc, leaving each summary statement hard to verify.
By Shuo Guan
The paper introduces Knowledge Pull Requests (KPRs), a framework that enables continual document authoring by making each change interpretable. KPRs extract claims from new knowledge sources, filter and route them to appropriate sections, and flag conflicts with existing content, producing a ChangeLog that separates knowledge changes from textual edits. Experiments on revising Wikipedia and updating query-driven reports show that KPRs integrate more information, better preserve existing content, and improve question answering performance compared to rewriting from scratch or using frontier models with search.
By Alexander Martin, Benjamin Van Durme
arXiv:2607.12441v3 Announce Type: replace
Abstract: Wikipedia plays a key role in shaping public understanding of science, and its openly accessible revision history is a unique record of how scienti...
By Omer Ehrlich, Nitzan Barzilay, Rona Aviram, Tom Hope
The paper introduces CAMS, a Claim‑Anchored Multi‑Document Summarization framework that decomposes source documents into atomic claims, resolves provenance deterministically from verbatim quotes to token spans, clusters equivalent claims across documents, and rewrites summaries so each sentence ends with claim identifiers linking back to source spans. CAMS separates provenance (an invariant for each emitted sentence) from faithfulness (an objective encouraged by selection, rewriting, and verification). Evaluations on MultiNews, DiverseSumm, and zero‑shot WCEP show that CAMS matches strong baselines in summary quality while improving faithfulness and citation precision, raising attribution accuracy from 38% to 64% and reducing human verification time per claim by 3.4×.
By Shuo Guan
The paper introduces an open, modular AI framework that automatically detects and structures evidence of social tipping points in climate literature at the passage level. It integrates a DistilBERT boundary splitter, an iteratively augmented RoBERTa classifier, a Mistral 7B rewrite model, a LLaMA 3.2 3B rating model, and a Milvus vector store, all accessible via a Streamlit interface. Evaluation on a GPT‑4.1‑labelled benchmark and expert‑reviewed set shows the splitter outperforms competitors and the RoBERTa detector achieves high accuracy and agreement.
By Kavindu Perera, Mohammad Abaeiani, Ekaterina Gilman, Lauri Loven, Mourad Oussalah, Tassos Kanellos, Beatrice Gobbo, Dante Adami, Nicol\`o Ferriani, Maximiliano Romero, Pierre Rossel, Marc Bonazountas, Christina Deligianni, Nikos Xyderis, Artur Bogucki, Lampros Argyriou, Prasasthy Balasubramanian
arXiv:2609.14770v1 Announce Type: cross
Abstract: Generalisations are common in scientific communication, even though they are semantically ambiguous. An automated method is needed to identify and ca...
By Chenxin Diao, Nataliya Stepanova, Emily Allaway
arXiv:2607. 28618v1 Announce Type: cross Abstract: Chemistry literature synthesis often requires assembling specific findings scattered across many publications, yet existing literature-search systems primarily return ranked document lists.
By Bing Yan, Gregory Wolfe, Stefano Martiniani, Kyunghyun Cho
The paper introduces CLAIMPROBE, a claim-level audit tool that breaks down deep-research reports into individual claims and evaluates them for hallucination, misattribution, citation hygiene, and necessary-fact recall against retrieved evidence. Using CLAIMPROBE, the authors show that even high-scoring deep-research pipelines can omit key evidence and misattribute claims. They also propose CLAIMWRITER, a hierarchical claim-based writer that extracts source facts, maps them to an outline, and drafts sections from a source-linked claim representation, which reduces hallucination by 2.6 to 4.5 times and improves necessary-fact recall by 1.2 to 1.7 times while preserving overall report quality and enabling efficient localized revisions.
By Hiroaki Hayashi, Pranav Narayanan Venkit, Prafulla Kumar Choubey, Chien-Sheng Wu
The paper investigates how large language models can extract contextualized data from scientific literature. It presents four workflows: expert‑written prompts, self‑generated prompts, autonomous literature discovery, and dataset creation from guidelines. While models perform well with prompts, they struggle with context, hallucinate references, and still need human oversight for final validation.
By Valentin Romanov, Monique Bax, Steven Niederer