arXiv Computation and Language

Peerify: Benchmarking Peer-Review Claim Verification

arXiv AI
Aug 5

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

arXiv:2608. 03581v1 Announce Type: cross Abstract: AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are.

By Alexander M. Fichtl, Lukas Ellinger, Josefin Kelber, Kry\v{s}tof Ol\'ik, Georg Groh
Hugging Face Trending Papers
Aug 4

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are. We address these questions by first surveying reviewer-facing AI policies across 111 leading AI/NLP conferences and medical journals, revealing substantial regulation differences between the two communities.

arXiv AI
Jun 24

Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable

arXiv:2603. 20450v2 Announce Type: replace-cross Abstract: A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polishing, paraphrasing, and grammar correction of otherwise human-written reviews.

By Rounak Saha, Gurusha Juneja, Dayita Chaudhuri, Naveeja Sajeevan, Nihar B Shah, Danish Pruthi
arXiv AI
Sep 4

HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews

HalluPeer is a new benchmark designed to detect hallucinations in scientific peer reviews. It provides aligned triples of paper content, human-written reviews, and hallucination-injected reviews, annotated for detection, classification, and localization. Experiments on 12K papers and 38K reviews show that existing detectors struggle to separate hallucinations from legitimate critique, and that HalluPeer-defined hallucination patterns occur in real peer reviews.

By Tzu-Ling Lin, Dong-Ting Yao, Teng-Fang Hsiao, Wei-Chih Chen, Hong-Han Shuai
Hugging Face Trending Papers
Sep 3

HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews

HalluPeer is a new benchmark designed to detect hallucinations in scientific peer reviews. It provides aligned triples of paper content, human-written reviews, and hallucination-injected reviews, annotated for detection, classification, and localization. Experiments on 12K papers and 38K reviews show that current detectors struggle to distinguish hallucinations from legitimate critique, and real peer reviews contain HalluPeer-defined hallucination patterns, underscoring the need for source-aware verification.

arXiv Machine Learning
Sep 21

When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

Large language models are increasingly used as automated reviewers in scientific evaluation, creating a recursive feedback loop where later reviewers learn from earlier model-generated judgments. A study using Llama 3.1 8B fine‑tuned on ICLR reviews shows that incorporating synthetic reviews compresses rating distributions and reduces semantic diversity, a phenomenon termed scientific‑judgment collapse. To counter this, the authors introduce TrustReviewer, an open‑source LLM system that curates training data and applies paired activation steering at test time to preserve judgment diversity and improve recommendation alignment.

By Sy-Tuyen Ho, Minghui Liu, Furong Huang