arXiv Computation and Language

Pair-Level Essay-Scale Republication and Reuse from Fragmented Historical Text Reuse: A Workflow Study on Eighteenth-Century Books and Newspapers

This study tackles the challenge of identifying essay‑scale republication and reuse from fragmented text evidence, focusing on David Hume essays in eighteenth‑century books and newspapers. It compares a staged rule‑based workflow, baseline decision‑tree and LLM approaches, and automated rule adaptation, finding that pair‑level feature aggregation achieves high F1 scores and that the final workflow offers the best precision‑recall balance. Manual audits confirm all predicted positives as genuine republications, demonstrating the method’s effectiveness in producing compact, auditable candidate sets for historical analysis.

arXiv Computation and Language
Sep 22

Re:CAP - Auditing Retrieval Coverage in Production RAG Pipelines

Re:CAP is a reference‑free audit loop for retrieval‑augmented generation (RAG) pipelines that probes for missing documents instead of enumerating all relevant ones. It identifies covered topics, generates probing questions, retrieves candidate documents, and uses an LLM judge to keep only those that add new information. On several benchmarks, Re:CAP recovers a significant portion of gold documents that flat BM25 or hybrid retrieval misses, and human evaluation shows most of these documents add new information.

By Aviral Joshi, Hanoz Bhathena, Max Nelson, Saket Sharma
arXiv AI
Jul 28

VecTree-RAG: An Agentic Retrieval-Augmented Generation Framework Combining Vector and Tree Retrieval for Efficiency and Accuracy

arXiv:2607. 23006v1 Announce Type: cross Abstract: Scientific question answering requires a retrieval system to solve two distinct problems: identifying which papers are relevant and locating the supporting evidence within those papers.

By Xinyan Zhong, Yuwei Shi, Yuqi Wei, Chen Shen, Tianhang Zhou, Zhenghao Wu
arXiv Machine Learning
Sep 11

Narrative Consolidation: Formulating a New Task for Unifying Multi-Perspective Accounts

The paper introduces Narrative Consolidation, a new NLP task that aims to merge overlapping narrative documents—such as legal testimonies or historical accounts—into a single, chronologically coherent text, rather than merely compressing them. It defines the task, proposes an evaluation framework, and presents the Gospel Consolidation Language Resource, a benchmark built from the four Biblical Gospels with 169 canonical events and cross‑document alignments. Experiments show that providing an explicit temporal backbone dramatically improves performance, a simple length heuristic outperforms graph‑based methods, and temporal edges are the key discriminative signal.

By Roger A. Finger, Eduardo G. Cortes, Sandro J. Rigo, Gabriel de O. Ramos
arXiv AI
Jul 1

RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora

arXiv:2604. 19047v2 Announce Type: replace-cross Abstract: Existing QA benchmarks typically assume distinct documents with minimal overlap, yet real-world retrieval-augmented generation (RAG) systems operate on corpora such as financial reports, legal codes, and patents, where information is highly redundant and documents exhibit strong inter-document similarity.

By Hanjun Cho, Jay-Yoon Lee
arXiv AI
Aug 26

Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes

The paper audits a 366‑day autobiographical book generated by a large language model (LLM) against an independent verification corpus. Using a four‑level rubric, 354 of the 366 days (96.7%) failed verification, with only 12 days containing corroborated scenes and 19 days containing actively contradicted claims. Regenerating the same days with current models yielded 100% verification failure, while grounding the generation in the subject’s own corpus improved the rate to 83.3% but still left substantial residual failure.

By Heather Renze
arXiv AI
Sep 4

More Criticism Does Not Make a Better Review: EquiReview-R

The paper introduces EquiReview‑R, an AI‑assisted review system that treats omission and over‑critique as distinct risks and refines a structured concern set using evidence‑linked reasoning. It demonstrates that more criticism does not guarantee a better review, showing that many high‑recall reviews lack definitive evidence for concerns and that revision before further search is essential. On a held‑out set of papers, EquiReview‑R meets non‑inferiority for major omission, cuts major over‑critique from 15.5 % to 8.1 %, and stops on 52.4 % of papers, with gains attributed to revision rather than extra inference.

By Zexing Zhang, Jichao Li, Tianyang Lei, Yude Fu, Yang Kewei