PaperDoctor is an agent framework that provides evidence‑grounded, actionable feedback for scientific papers before submission. It evaluates writing, layout, references, code, theory, prior work, and experiments through a three‑layer hierarchical system, linking each critique to specific evidence and revision suggestions. The system selectively rebuilds and reruns experiments to uncover reproducibility gaps, and an interactive interface lets authors explore findings tied to their manuscript.
By Kevin Qinghong Lin, Siyuan Hu, Pan Lu, Yu Chen, Yanzhe Chen, Owen Queen, Yupeng Chen, Jialin Yu, Junchi Yu, Zifeng Ding, Yuanfeng Ji, Sheng Liu, Jindong Gu, Linjie Li, Mike Zheng Shou, Philip Torr, James Zou
arXiv:2609.14738v1 Announce Type: new
Abstract: Automated reviewing systems are increasingly evaluated based on the quality of the reviews they produce. Yet a review is only useful if acting on it le...
By Vidushee Vats, Karun Sharma, Shengzhi Li, Shichao Pei
The paper presents an LLM-driven framework that splits peer reviews into argumentative segments, detects multiple co-occurring issues such as lazy thinking and lack of specificity, and generates targeted, guideline-aware feedback using issue-specific templates. An iterative, reranking-based generation algorithm refines the feedback, and a controlled rewriting study shows it can reduce guideline violations by up to 92.4%. The authors also release LazyReviewPlus, a multi-label dataset of 1,309 sentences annotated for detecting lazy thinking and lack of specificity.
By Sukannya Purkayastha, Qile Wan, Anne Lauscher, Lizhen Qu, Iryna Gurevych
arXiv:2609.05947v1 Announce Type: new
Abstract: Peer review is central to quality control in science. However, existing evaluations of AI-assisted peer review mainly focus on the overall quality of g...
By Siming Yuan, Xueyi Zhang, Wangze Ni, Tianfang Xiao, Shimin Di, Jia Zhu, Zhuoren Jiang, Rong Tan, Lei Chen, Kui Ren
arXiv:2601. 14171v2 Announce Type: replace Abstract: Writing effective rebuttals is a high-stakes task that demands more than linguistic fluency, as it requires precise alignment between reviewer intent and manuscript details.
By Qianli Ma, Chang Guo, Zhiheng Tian, Siyu Wang, Jipeng Xiao, Yuanhao Yue, Zhipeng Zhang
arXiv:2608. 08975v1 Announce Type: cross Abstract: As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions.
By Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou
arXiv:2606. 06025v1 Announce Type: cross Abstract: Scientific peer review generation has attracted increasing attention for reducing reviewing burdens and providing timely feedback.
By Xinpeng Qiu, Wang Yihu, Zhifeng Liu, Xiaochen Wang, Jimin Wang
arXiv:2607. 26066v1 Announce Type: cross Abstract: The growing volume of scientific submissions has motivated interest in using large language models (LLMs) to assist peer review.
By Ranjitha Shivaprasad Ballakuraya, Arash Mahyari, Ashok Srinivasan
arXiv:2608.20488v1 Announce Type: new
Abstract: AI tools increasingly support tasks across the scientific research cycle, from experiment design and manuscript preparation to peer review. At the same...
By Anirudh Sundar, Min Chen, Divya Tadimeti, Gemma Zhang, Alice Li, Nigel Boachie Kumankumah, Pavan Uttej Ravva, Sadid Hasan, Somya Chatterjee, Pruthvi Prakash Navada, Xiao Wang, Yue Kang, Sulaiman Vesal, Larry Heck
Tree-of-Concerns is a multi‑agent framework that uses specialized skeptic personas to conduct parallel debate trees, each focusing on a specific category of potential limitations in scientific papers. The system employs structured, evidence‑grounded argumentation and a panel review mechanism to correct drift and miscalibration, ultimately extracting unstated limitations. Experiments on the ToC‑Bench benchmark show that the approach improves precision by 79% and coverage by 11% over leading baselines, providing reviewers with specific, evidence‑based concerns for systematic evaluation.
By Sahil Mishra, Niranjan Rajeev, Tanmoy Chakraborty
arXiv:2606. 31478v1 Announce Type: new Abstract: Autonomous research agents can now draft hypotheses, write code, run experiments, and produce papers, but they remain brittle when experiments fail.
By Jie Ma, Binfei Chu, Jie Gao, Jinlu Zhang, Yiwei Ma, Yi Tan, Jiayi Ji, Xiaoshuai Sun, Rongrong Ji
arXiv:2602. 18446v2 Announce Type: replace-cross Abstract: Users increasingly rely on Large Language Models (LLMs) for Deep Research, using them to synthesize diverse sources into structured reports that support understanding and action.
By Jujia Zhao, Zhaoxin Huan, Zihan Wang, Xiaolu Zhang, Jun Zhou, Suzan Verberne, Zhaochun Ren