arXiv AI By Ling Yue, Chaoqian Ouyang, Hang Xu, Ruijun Huang, Yuchen Liu, Libin Zheng, Wei Liu, Shaowu Pan, Shimin Di, Min-Ling Zhang

FactReview: Evidence-Grounded Peer Review with Execution-Based Claim Verification

Read the original on arXiv AI →

arXiv:2604. 04074v4 Announce Type: replace Abstract: Large language model (LLM)-based reviewing systems typically assess manuscripts in isolation, leaving literature- and code-dependent claims difficult to verify.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 23

Peerify: Benchmarking Peer-Review Claim Verification

arXiv:2609.25046v1 Announce Type: new Abstract: Peer review plays a central role in scholarly publishing, yet verifying whether reviewer claims are supported by manuscript evidence remains a largely...

By Alireza Daghighfarsoodeh, Sajad Ebrahimi, Ali Ghorbanpour, Soroush Sadeghian, Radin Cheraghi, Negar Arabzadeh, Ebrahim Bagheri
arXiv AI
Sep 1

Redesigning and Auditing Deep Research Writing for Faithful Reports

The paper introduces CLAIMPROBE, a claim-level audit tool that breaks down deep-research reports into individual claims and evaluates them for hallucination, misattribution, citation hygiene, and necessary-fact recall against retrieved evidence. Using CLAIMPROBE, the authors show that even high-scoring deep-research pipelines can omit key evidence and misattribute claims. They also propose CLAIMWRITER, a hierarchical claim-based writer that extracts source facts, maps them to an outline, and drafts sections from a source-linked claim representation, which reduces hallucination by 2.6 to 4.5 times and improves necessary-fact recall by 1.2 to 1.7 times while preserving overall report quality and enabling efficient localized revisions.

By Hiroaki Hayashi, Pranav Narayanan Venkit, Prafulla Kumar Choubey, Chien-Sheng Wu
arXiv AI
Sep 10

More Than Mimicking Reviewers: Evaluating LLMs for Pre-Submission Peer Review

The paper introduces an author-facing large language model (LLM) system that generates a broad set of atomic concerns about a manuscript and compresses them into a concise report, aiming to provide early peer‑review feedback. Evaluations on 3,398 ICLR 2026 submissions show that the system covers 44.9% of historical reviewer issues on a diagnostic set, rising to 78.7% strict coverage and 84.9% seriousness‑weighted coverage after deduplication and refill, using 3.6× more requests and 5.2× more tokens. Ablation studies reveal that representative selection and matcher sensitivity are key factors limiting the compression quality.

By Pouya Parsa, Amin Rezaei