arXiv Computation and Language By Ke Shu, Kira Hinderks, Eetu M\"akel\"a, Mikko Tolonen

Pair-Level Essay-Scale Republication and Reuse from Fragmented Historical Text Reuse: A Workflow Study on Eighteenth-Century Books and Newspapers

Read the original on arXiv Computation and Language →

This study tackles the challenge of identifying essay‑scale republication and reuse from fragmented text evidence, focusing on David Hume essays in eighteenth‑century books and newspapers. It compares a staged rule‑based workflow, baseline decision‑tree and LLM approaches, and automated rule adaptation, finding that pair‑level feature aggregation achieves high F1 scores and that the final workflow offers the best precision‑recall balance. Manual audits confirm all predicted positives as genuine republications, demonstrating the method’s effectiveness in producing compact, auditable candidate sets for historical analysis.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 22

Re:CAP - Auditing Retrieval Coverage in Production RAG Pipelines

Re:CAP is a reference‑free audit loop for retrieval‑augmented generation (RAG) pipelines that probes for missing documents instead of enumerating all relevant ones. It identifies covered topics, generates probing questions, retrieves candidate documents, and uses an LLM judge to keep only those that add new information. On several benchmarks, Re:CAP recovers a significant portion of gold documents that flat BM25 or hybrid retrieval misses, and human evaluation shows most of these documents add new information.

By Aviral Joshi, Hanoz Bhathena, Max Nelson, Saket Sharma
arXiv AI
Jul 28

VecTree-RAG: An Agentic Retrieval-Augmented Generation Framework Combining Vector and Tree Retrieval for Efficiency and Accuracy

arXiv:2607. 23006v1 Announce Type: cross Abstract: Scientific question answering requires a retrieval system to solve two distinct problems: identifying which papers are relevant and locating the supporting evidence within those papers.

By Xinyan Zhong, Yuwei Shi, Yuqi Wei, Chen Shen, Tianhang Zhou, Zhenghao Wu
arXiv Machine Learning
Sep 11

Narrative Consolidation: Formulating a New Task for Unifying Multi-Perspective Accounts

The paper introduces Narrative Consolidation, a new NLP task that aims to merge overlapping narrative documents—such as legal testimonies or historical accounts—into a single, chronologically coherent text, rather than merely compressing them. It defines the task, proposes an evaluation framework, and presents the Gospel Consolidation Language Resource, a benchmark built from the four Biblical Gospels with 169 canonical events and cross‑document alignments. Experiments show that providing an explicit temporal backbone dramatically improves performance, a simple length heuristic outperforms graph‑based methods, and temporal edges are the key discriminative signal.

By Roger A. Finger, Eduardo G. Cortes, Sandro J. Rigo, Gabriel de O. Ramos