arXiv AI

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

arXiv:2608. 03581v1 Announce Type: cross Abstract: AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are.

Hugging Face Trending Papers
Aug 4

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are. We address these questions by first surveying reviewer-facing AI policies across 111 leading AI/NLP conferences and medical journals, revealing substantial regulation differences between the two communities.

arXiv AI
Jun 24

Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable

arXiv:2603. 20450v2 Announce Type: replace-cross Abstract: A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polishing, paraphrasing, and grammar correction of otherwise human-written reviews.

By Rounak Saha, Gurusha Juneja, Dayita Chaudhuri, Naveeja Sajeevan, Nihar B Shah, Danish Pruthi
arXiv AI
Aug 28

FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes

FIRSTPASS is a large-scale peer review dataset that captures complete multi-round editorial dialogues from a multidisciplinary high-impact journal, Nature Communications. It contains 3,668 records across five scientific domains—biology, chemistry, neuroscience, physics, and earth science—encompassing initial referee reports, author responses, and updated reviewer assessments. Each record is labeled with an outcome (STANDARD or EXTENDED) based on editorial decisions, and the dataset includes detailed parsing pipelines and evaluation scripts for reproducible AI benchmarking.

By Prabhjot Singh, Somnath Luitel, Manmeet Singh, Josh Durkee
arXiv Machine Learning
Sep 21

When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

Large language models are increasingly used as automated reviewers in scientific evaluation, creating a recursive feedback loop where later reviewers learn from earlier model-generated judgments. A study using Llama 3.1 8B fine‑tuned on ICLR reviews shows that incorporating synthetic reviews compresses rating distributions and reduces semantic diversity, a phenomenon termed scientific‑judgment collapse. To counter this, the authors introduce TrustReviewer, an open‑source LLM system that curates training data and applies paired activation steering at test time to preserve judgment diversity and improve recommendation alignment.

By Sy-Tuyen Ho, Minghui Liu, Furong Huang
arXiv AI
Aug 5

How Closely Do LLM Reviews Align with Human Peer Review?

arXiv:2608. 03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting.

By Abraham Camelo-Guerrero, Jairo Diaz-Rodriguez