arXiv AI

Redesigning and Auditing Deep Research Writing for Faithful Reports

The paper introduces CLAIMPROBE, a claim-level audit tool that breaks down deep-research reports into individual claims and evaluates them for hallucination, misattribution, citation hygiene, and necessary-fact recall against retrieved evidence. Using CLAIMPROBE, the authors show that even high-scoring deep-research pipelines can omit key evidence and misattribute claims. They also propose CLAIMWRITER, a hierarchical claim-based writer that extracts source facts, maps them to an outline, and drafts sections from a source-linked claim representation, which reduces hallucination by 2.6 to 4.5 times and improves necessary-fact recall by 1.2 to 1.7 times while preserving overall report quality and enabling efficient localized revisions.

arXiv AI
Sep 25

Claim-Gated Source-Risk Auditing for Generative Search

The paper introduces a claim‑gated audit framework for generative search, ensuring that a query, source, and answer tuple is only considered resolved when relationship evidence, answer adoption, materiality, and disclosure are all present. It distinguishes this audit endpoint from citation support and review priority, tying decisions to versioned evidence spans and implementing a reference checker to enforce the contract. Experiments on a synthetic dataset confirm that the system correctly handles all 81 predicate combinations and rejects 192 malformed records, while ablation studies isolate endpoint logic from missing‑evidence handling.

By Kainan Zhou, Chuhong Xu, Gangzhen Qian, Zhaoyi Li
arXiv Computation and Language
Sep 23

Peerify: Benchmarking Peer-Review Claim Verification

arXiv:2609.25046v1 Announce Type: new Abstract: Peer review plays a central role in scholarly publishing, yet verifying whether reviewer claims are supported by manuscript evidence remains a largely...

By Alireza Daghighfarsoodeh, Sajad Ebrahimi, Ali Ghorbanpour, Soroush Sadeghian, Radin Cheraghi, Negar Arabzadeh, Ebrahim Bagheri
arXiv AI
Aug 20

DeepWeaver: Bridging the Evidence Synthesis Gap in Open-Ended Question Answering

DeepWeaver is a framework designed to improve open‑ended question answering by weaving noisy retrieved evidence into comprehensive, well‑cited answers. It introduces Thought Block Chains (TBCs) that organize claims, key information, and supporting evidence, and uses subordinate TBCs to refine and expand the evidence before final generation. Evaluations on LoQA and DeepResearch Bench show that DeepWeaver enhances content sufficiency, citation grounding, and detail preservation across multiple LLMs.

By Xujia Wang, Yizhe Zhang, Bin Xu, Lei Hou, Juanzi Li
arXiv AI
Jun 2

Med-V1: Small Language Models for Zero-shot and Scalable Biomedical Evidence Attribution

arXiv:2603. 05308v3 Announce Type: replace-cross Abstract: Assessing whether an article supports an assertion is essential for hallucination detection and claim verification.

By Qiao Jin, Yin Fang, Lauren He, Yifan Yang, Guangzhi Xiong, Zhizheng Wang, Nicholas Wan, Joey Chan, Donald C. Comeau, Robert Leaman, Charalampos S. Floudas, Aidong Zhang, Michael F. Chiang, Yifan Peng, Zhiyong Lu
Hugging Face Trending Papers
Aug 19

DeepWeaver: Bridging the Evidence Synthesis Gap in Open-Ended Question Answering

DeepWeaver addresses the evidence synthesis gap in open‑ended question answering by weaving noisy retrieved evidence into comprehensive answers. It introduces Thought Block Chains (TBCs) that organize claims, key information, and citations, allowing the system to revise and expand evidence before final generation. Evaluations on LoQA and DeepResearch Bench show improved content sufficiency, citation grounding, and detail preservation across multiple LLMs.

arXiv AI
Sep 7

Attributable by Construction: Claim-Anchored Provenance for Multi-Document Summarization

The paper introduces CAMS, a Claim‑Anchored Multi‑Document Summarization framework that decomposes source documents into atomic claims, resolves provenance deterministically from verbatim quotes to token spans, clusters equivalent claims across documents, and rewrites summaries so each sentence ends with claim identifiers linking back to source spans. CAMS separates provenance (an invariant for each emitted sentence) from faithfulness (an objective encouraged by selection, rewriting, and verification). Evaluations on MultiNews, DiverseSumm, and zero‑shot WCEP show that CAMS matches strong baselines in summary quality while improving faithfulness and citation precision, raising attribution accuracy from 38% to 64% and reducing human verification time per claim by 3.4×.

By Shuo Guan