arXiv AI By Zexing Zhang, Jichao Li, Tianyang Lei, Yude Fu, Yang Kewei

More Criticism Does Not Make a Better Review: EquiReview-R

Read the original on arXiv AI →

The paper introduces EquiReview‑R, an AI‑assisted review system that treats omission and over‑critique as distinct risks and refines a structured concern set using evidence‑linked reasoning. It demonstrates that more criticism does not guarantee a better review, showing that many high‑recall reviews lack definitive evidence for concerns and that revision before further search is essential. On a held‑out set of papers, EquiReview‑R meets non‑inferiority for major omission, cuts major over‑critique from 15.5 % to 8.1 %, and stops on 52.4 % of papers, with gains attributed to revision rather than extra inference.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
3d ago

PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress

PaperDoctor is an agent framework that provides evidence‑grounded, actionable feedback for scientific papers before submission. It evaluates writing, layout, references, code, theory, prior work, and experiments through a three‑layer hierarchical system, linking each critique to specific evidence and revision suggestions. The system selectively rebuilds and reruns experiments to uncover reproducibility gaps, and an interactive interface lets authors explore findings tied to their manuscript.

By Kevin Qinghong Lin, Siyuan Hu, Pan Lu, Yu Chen, Yanzhe Chen, Owen Queen, Yupeng Chen, Jialin Yu, Junchi Yu, Zifeng Ding, Yuanfeng Ji, Sheng Liu, Jindong Gu, Linjie Li, Mike Zheng Shou, Philip Torr, James Zou
arXiv Computation and Language
Sep 3

Reviewing the Reviewer: LLM-Assisted Reviewer Feedback Generation for Guideline Compliance

The paper presents an LLM-driven framework that splits peer reviews into argumentative segments, detects multiple co-occurring issues such as lazy thinking and lack of specificity, and generates targeted, guideline-aware feedback using issue-specific templates. An iterative, reranking-based generation algorithm refines the feedback, and a controlled rewriting study shows it can reduce guideline violations by up to 92.4%. The authors also release LazyReviewPlus, a multi-label dataset of 1,309 sentences annotated for detecting lazy thinking and lack of specificity.

By Sukannya Purkayastha, Qile Wan, Anne Lauscher, Lizhen Qu, Iryna Gurevych
arXiv AI
Aug 11

How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review

arXiv:2608. 08975v1 Announce Type: cross Abstract: As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions.

By Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou