arXiv AI
Sep 10

More Than Mimicking Reviewers: Evaluating LLMs for Pre-Submission Peer Review

The paper introduces an author-facing large language model (LLM) system that generates a broad set of atomic concerns about a manuscript and compresses them into a concise report, aiming to provide early peer‑review feedback. Evaluations on 3,398 ICLR 2026 submissions show that the system covers 44.9% of historical reviewer issues on a diagnostic set, rising to 78.7% strict coverage and 84.9% seriousness‑weighted coverage after deduplication and refill, using 3.6× more requests and 5.2× more tokens. Ablation studies reveal that representative selection and matcher sensitivity are key factors limiting the compression quality.

By Pouya Parsa, Amin Rezaei
arXiv AI
Jun 24

Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable

arXiv:2603. 20450v2 Announce Type: replace-cross Abstract: A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polishing, paraphrasing, and grammar correction of otherwise human-written reviews.

By Rounak Saha, Gurusha Juneja, Dayita Chaudhuri, Naveeja Sajeevan, Nihar B Shah, Danish Pruthi
Hugging Face Trending Papers
Aug 4

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are. We address these questions by first surveying reviewer-facing AI policies across 111 leading AI/NLP conferences and medical journals, revealing substantial regulation differences between the two communities.

arXiv AI
Aug 5

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

arXiv:2608. 03581v1 Announce Type: cross Abstract: AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are.

By Alexander M. Fichtl, Lukas Ellinger, Josefin Kelber, Kry\v{s}tof Ol\'ik, Georg Groh
arXiv Computation and Language
Sep 24

How Much Were You Told? Measuring External Information in Peer Reviews

The paper introduces Self‑Conditioning, an unsupervised, information‑theoretic estimator that measures the amount of external information in peer reviews. It compares the likelihood of a review under its original production context with the likelihood when that context is augmented by hints extracted from the review itself. On the IntelLabs benchmark, Self‑Conditioning can perfectly distinguish fully‑delegated reviews from machine‑polished ones, remains largely insensitive to surface rewriting, and shows that increased external input drives scores toward human‑like values, unlike standard ATD baselines.

By Matthieu Dubois, Pablo Piantanida, Fran\c{c}ois Yvon