arXiv AI

Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review

arXiv:2607. 22553v1 Announce Type: cross Abstract: Peer review is an essential process in scientific research, yet the growing workload has made its automation increasingly necessary.

Hugging Face Trending Papers
Aug 4

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are. We address these questions by first surveying reviewer-facing AI policies across 111 leading AI/NLP conferences and medical journals, revealing substantial regulation differences between the two communities.

arXiv AI
Aug 5

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

arXiv:2608. 03581v1 Announce Type: cross Abstract: AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are.

By Alexander M. Fichtl, Lukas Ellinger, Josefin Kelber, Kry\v{s}tof Ol\'ik, Georg Groh
arXiv AI
Aug 5

How Closely Do LLM Reviews Align with Human Peer Review?

arXiv:2608. 03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting.

By Abraham Camelo-Guerrero, Jairo Diaz-Rodriguez
arXiv Computation and Language
Sep 3

Reviewing the Reviewer: LLM-Assisted Reviewer Feedback Generation for Guideline Compliance

The paper presents an LLM-driven framework that splits peer reviews into argumentative segments, detects multiple co-occurring issues such as lazy thinking and lack of specificity, and generates targeted, guideline-aware feedback using issue-specific templates. An iterative, reranking-based generation algorithm refines the feedback, and a controlled rewriting study shows it can reduce guideline violations by up to 92.4%. The authors also release LazyReviewPlus, a multi-label dataset of 1,309 sentences annotated for detecting lazy thinking and lack of specificity.

By Sukannya Purkayastha, Qile Wan, Anne Lauscher, Lizhen Qu, Iryna Gurevych
arXiv AI
Jun 24

Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable

arXiv:2603. 20450v2 Announce Type: replace-cross Abstract: A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polishing, paraphrasing, and grammar correction of otherwise human-written reviews.

By Rounak Saha, Gurusha Juneja, Dayita Chaudhuri, Naveeja Sajeevan, Nihar B Shah, Danish Pruthi
arXiv AI
Sep 4

HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews

HalluPeer is a new benchmark designed to detect hallucinations in scientific peer reviews. It provides aligned triples of paper content, human-written reviews, and hallucination-injected reviews, annotated for detection, classification, and localization. Experiments on 12K papers and 38K reviews show that existing detectors struggle to separate hallucinations from legitimate critique, and that HalluPeer-defined hallucination patterns occur in real peer reviews.

By Tzu-Ling Lin, Dong-Ting Yao, Teng-Fang Hsiao, Wei-Chih Chen, Hong-Han Shuai
arXiv Computation and Language
Sep 24

How Much Were You Told? Measuring External Information in Peer Reviews

The paper introduces Self‑Conditioning, an unsupervised, information‑theoretic estimator that measures the amount of external information in peer reviews. It compares the likelihood of a review under its original production context with the likelihood when that context is augmented by hints extracted from the review itself. On the IntelLabs benchmark, Self‑Conditioning can perfectly distinguish fully‑delegated reviews from machine‑polished ones, remains largely insensitive to surface rewriting, and shows that increased external input drives scores toward human‑like values, unlike standard ATD baselines.

By Matthieu Dubois, Pablo Piantanida, Fran\c{c}ois Yvon