arXiv Computation and Language

Drift Inspector: Exploring and Measuring Scientific Drift with Atomic Contribution Claims

Drift Inspector is an open‑source system that extracts Atomic Contribution Claims (ACCs) from scientific abstracts using an LLM, then clusters these claims over time to map how a research field evolves. Applied to six years of EMNLP, the tool reveals a shift from classic NLP tasks toward LLM‑era capabilities such as reasoning and multimodality—trends that keyword or whole‑abstract counts miss. The pipeline has also processed the entire ACL Anthology, yielding 346,000 claims from 80,000 abstracts across 423 venues, with human‑validated extraction and clustering aligned to an external taxonomy.

arXiv AI
Sep 7

Attributable by Construction: Claim-Anchored Provenance for Multi-Document Summarization

The paper introduces CAMS, a Claim‑Anchored Multi‑Document Summarization framework that decomposes source documents into atomic claims, resolves provenance deterministically from verbatim quotes to token spans, clusters equivalent claims across documents, and rewrites summaries so each sentence ends with claim identifiers linking back to source spans. CAMS separates provenance (an invariant for each emitted sentence) from faithfulness (an objective encouraged by selection, rewriting, and verification). Evaluations on MultiNews, DiverseSumm, and zero‑shot WCEP show that CAMS matches strong baselines in summary quality while improving faithfulness and citation precision, raising attribution accuracy from 38% to 64% and reducing human verification time per claim by 3.4×.

By Shuo Guan
arXiv AI
Sep 1

Redesigning and Auditing Deep Research Writing for Faithful Reports

The paper introduces CLAIMPROBE, a claim-level audit tool that breaks down deep-research reports into individual claims and evaluates them for hallucination, misattribution, citation hygiene, and necessary-fact recall against retrieved evidence. Using CLAIMPROBE, the authors show that even high-scoring deep-research pipelines can omit key evidence and misattribute claims. They also propose CLAIMWRITER, a hierarchical claim-based writer that extracts source facts, maps them to an outline, and drafts sections from a source-linked claim representation, which reduces hallucination by 2.6 to 4.5 times and improves necessary-fact recall by 1.2 to 1.7 times while preserving overall report quality and enabling efficient localized revisions.

By Hiroaki Hayashi, Pranav Narayanan Venkit, Prafulla Kumar Choubey, Chien-Sheng Wu
arXiv Computation and Language
Sep 23

Peerify: Benchmarking Peer-Review Claim Verification

arXiv:2609.25046v1 Announce Type: new Abstract: Peer review plays a central role in scholarly publishing, yet verifying whether reviewer claims are supported by manuscript evidence remains a largely...

By Alireza Daghighfarsoodeh, Sajad Ebrahimi, Ali Ghorbanpour, Soroush Sadeghian, Radin Cheraghi, Negar Arabzadeh, Ebrahim Bagheri
arXiv Computation and Language
Sep 23

Knowledge Pull Requests for Continual Document Authoring

The paper introduces Knowledge Pull Requests (KPRs), a framework that enables continual document authoring by making each change interpretable. KPRs extract claims from new knowledge sources, filter and route them to appropriate sections, and flag conflicts with existing content, producing a ChangeLog that separates knowledge changes from textual edits. Experiments on revising Wikipedia and updating query-driven reports show that KPRs integrate more information, better preserve existing content, and improve question answering performance compared to rewriting from scratch or using frontier models with search.

By Alexander Martin, Benjamin Van Durme
arXiv AI
2d ago

Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers

The paper "Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers" introduces SciSlopBench, a dataset of 390 AI‑generated papers paired with human‑written counterparts, and defines six measures across Structure, Argument, and Artifacts to detect scientific slop. The authors show that these measures can identify AI papers with 85.9% accuracy and that higher slop correlates with lower ICLR ratings and distinguishes rejected from accepted papers. They also propose SciSlopHarness, a framework that guides a fixed LLM to revise only evidence‑supported sections, reducing the AI‑human gap by 63% without human reference targets.

By Yerim Oh, Young-Jun Lee, Jaewoo Ahn, Gunhee Kim, Dongyeop Kang