AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are. We address these questions by first surveying reviewer-facing AI policies across 111 leading AI/NLP conferences and medical journals, revealing substantial regulation differences between the two communities.
arXiv:2608. 03581v1 Announce Type: cross Abstract: AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are.
By Alexander M. Fichtl, Lukas Ellinger, Josefin Kelber, Kry\v{s}tof Ol\'ik, Georg Groh
arXiv:2605. 03202v2 Announce Type: replace Abstract: Large language models offer a tempting solution to address the peer review crisis.
By Joachim Baumann, Jiaxin Pei, Sanmi Koyejo, Dirk Hovy
arXiv:2608. 03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting.
By Abraham Camelo-Guerrero, Jairo Diaz-Rodriguez
arXiv:2609.23264v1 Announce Type: new
Abstract: Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high sc...
By Shakiba Amirshahi, Sajad Ebrahimi, Hai Son Le, Negar Arabzadeh, Ebrahim Bagheri
The paper presents an LLM-driven framework that splits peer reviews into argumentative segments, detects multiple co-occurring issues such as lazy thinking and lack of specificity, and generates targeted, guideline-aware feedback using issue-specific templates. An iterative, reranking-based generation algorithm refines the feedback, and a controlled rewriting study shows it can reduce guideline violations by up to 92.4%. The authors also release LazyReviewPlus, a multi-label dataset of 1,309 sentences annotated for detecting lazy thinking and lack of specificity.
By Sukannya Purkayastha, Qile Wan, Anne Lauscher, Lizhen Qu, Iryna Gurevych