arXiv AI

Faithful by Construction: Claim-Anchored Attribution for Multi-Document Summarization

arXiv:2606. 23989v1 Announce Type: cross Abstract: End-to-end large language models (LLMs) produce fluent multi-document summaries but remain prone to hallucination, and the attributions they offer are typically coarse (whole documents or passages) and generated post hoc, leaving each summary statement hard to verify.

arXiv AI
Sep 7

Attributable by Construction: Claim-Anchored Provenance for Multi-Document Summarization

The paper introduces CAMS, a Claim‑Anchored Multi‑Document Summarization framework that decomposes source documents into atomic claims, resolves provenance deterministically from verbatim quotes to token spans, clusters equivalent claims across documents, and rewrites summaries so each sentence ends with claim identifiers linking back to source spans. CAMS separates provenance (an invariant for each emitted sentence) from faithfulness (an objective encouraged by selection, rewriting, and verification). Evaluations on MultiNews, DiverseSumm, and zero‑shot WCEP show that CAMS matches strong baselines in summary quality while improving faithfulness and citation precision, raising attribution accuracy from 38% to 64% and reducing human verification time per claim by 3.4×.

By Shuo Guan
arXiv Computation and Language
Sep 18

Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation

The paper investigates how the granularity of citations—sentence, paragraph, or document level—affects the performance of large language models in attributed generation tasks. Across models ranging from 8B to 120B parameters, enforcing fine‑grained, sentence‑level citations consistently reduces performance, with median losses of 40% and up to 338% on specific tasks, while overall answer correctness remains largely unchanged. The study finds that attribution quality peaks at intermediate, paragraph‑level granularity, suggesting that overly fine citations break semantic dependencies and overly coarse ones add noise, and that the optimal granularity depends on model scale and the amount of evidence required. whyItMatters":"The findings reveal that the conventional preference for fine‑grained citations can actually harm model performance, indicating that attribution standards should be tailored to the model’s semantic scope rather than fixed by convention."

By Hexuan Wang, Jingyu Zhang, Benjamin Van Durme, Daniel Khashabi
arXiv AI
Aug 20

DeepWeaver: Bridging the Evidence Synthesis Gap in Open-Ended Question Answering

DeepWeaver is a framework designed to improve open‑ended question answering by weaving noisy retrieved evidence into comprehensive, well‑cited answers. It introduces Thought Block Chains (TBCs) that organize claims, key information, and supporting evidence, and uses subordinate TBCs to refine and expand the evidence before final generation. Evaluations on LoQA and DeepResearch Bench show that DeepWeaver enhances content sufficiency, citation grounding, and detail preservation across multiple LLMs.

By Xujia Wang, Yizhe Zhang, Bin Xu, Lei Hou, Juanzi Li
arXiv Computation and Language
3d ago

Drift Inspector: Exploring and Measuring Scientific Drift with Atomic Contribution Claims

Drift Inspector is an open‑source system that extracts Atomic Contribution Claims (ACCs) from scientific abstracts using an LLM, then clusters these claims over time to map how a research field evolves. Applied to six years of EMNLP, the tool reveals a shift from classic NLP tasks toward LLM‑era capabilities such as reasoning and multimodality—trends that keyword or whole‑abstract counts miss. The pipeline has also processed the entire ACL Anthology, yielding 346,000 claims from 80,000 abstracts across 423 venues, with human‑validated extraction and clustering aligned to an external taxonomy.

By Vsevolod Karimov, Stepan Ostarkov, Anastasia Poroshina, Anatoly Frolov, Alexander Panchenko
Hugging Face Trending Papers
Aug 19

DeepWeaver: Bridging the Evidence Synthesis Gap in Open-Ended Question Answering

DeepWeaver addresses the evidence synthesis gap in open‑ended question answering by weaving noisy retrieved evidence into comprehensive answers. It introduces Thought Block Chains (TBCs) that organize claims, key information, and citations, allowing the system to revise and expand evidence before final generation. Evaluations on LoQA and DeepResearch Bench show improved content sufficiency, citation grounding, and detail preservation across multiple LLMs.

arXiv Computation and Language
Aug 27

Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization

The paper introduces Gavel, a framework for evaluating large language models (LLMs) on long-context legal summarization tasks. Gavel includes a reference-based component (Gavel-Ref) with checklist, residual-fact, and writing-style checks, and a reference-free component (Gavel-Agent) that assesses factual coverage directly from source documents. Experiments on 12 frontier LLMs reveal that models tend to omit key information more than hallucinate, perform well on simple checklist items but struggle with rare, complex items, and their performance degrades with longer cases. Gavel-Agent cuts token usage by at least 36% compared to traditional methods while maintaining competitive accuracy, and it also generalizes effectively to the medical domain.

By Yao Dou, Benjamin Mamut, Wei Xu