arXiv Machine Learning

FMMD: A multimodal multidisciplinary dataset of open peer reviews from F1000Research

FMMD is a multimodal, multidisciplinary dataset of open peer reviews from F1000Research that pairs manuscript-level visual and structural data with version‑specific reviewer reports and editorial decisions. It addresses key gaps in existing datasets by preserving precise alignment between review comments and the exact manuscript version, and by including a wide range of scientific disciplines beyond computer science. The dataset supports tasks such as visual‑semantic consistency classification, figure‑related review comment generation, and editorial decision prediction, providing a comprehensive empirical resource for multimodal automated scholarly paper review research.

arXiv AI
Aug 5

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

arXiv:2608. 03581v1 Announce Type: cross Abstract: AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are.

By Alexander M. Fichtl, Lukas Ellinger, Josefin Kelber, Kry\v{s}tof Ol\'ik, Georg Groh
Hugging Face Trending Papers
Aug 4

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are. We address these questions by first surveying reviewer-facing AI policies across 111 leading AI/NLP conferences and medical journals, revealing substantial regulation differences between the two communities.

arXiv AI
Jul 31

SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence

arXiv:2607. 27066v1 Announce Type: cross Abstract: Scientific figure assessment in peer review differs fundamentally from general image quality evaluation: a figure must be visually legible, faithfully support the manuscript's claims, and communicate evidence with a clear visual hierarchy.

By Chuanzhi Xu, Zihan Deng, Huiqi Liang, Chengkun Yue, Zhanlin Cui, Pengfei Ye, Weidong Cai
arXiv Computation and Language
Sep 16

PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress

PaperDoctor is an agent framework that provides evidence‑grounded, actionable feedback for scientific papers before submission. It evaluates writing, layout, references, code, theory, prior work, and experiments through a three‑layer hierarchical system, linking each critique to specific evidence and revision suggestions. The system selectively rebuilds and reruns experiments to uncover reproducibility gaps, and an interactive interface lets authors explore findings tied to their manuscript.

By Kevin Qinghong Lin, Siyuan Hu, Pan Lu, Yu Chen, Yanzhe Chen, Owen Queen, Yupeng Chen, Jialin Yu, Junchi Yu, Zifeng Ding, Yuanfeng Ji, Sheng Liu, Jindong Gu, Linjie Li, Mike Zheng Shou, Philip Torr, James Zou
arXiv AI
Aug 28

FIRSTPASS: A Multi-Domain, Multi-Round Peer Review Dataset Grounded in Real Editorial Outcomes

FIRSTPASS is a large-scale peer review dataset that captures complete multi-round editorial dialogues from a multidisciplinary high-impact journal, Nature Communications. It contains 3,668 records across five scientific domains—biology, chemistry, neuroscience, physics, and earth science—encompassing initial referee reports, author responses, and updated reviewer assessments. Each record is labeled with an outcome (STANDARD or EXTENDED) based on editorial decisions, and the dataset includes detailed parsing pipelines and evaluation scripts for reproducible AI benchmarking.

By Prabhjot Singh, Somnath Luitel, Manmeet Singh, Josh Durkee
arXiv Computation and Language
Sep 11

ReGround: Grounding Reviewer Comments in Multimodal Evidence

ReGround is a new large‑scale dataset that links 10,267 reviewer comments to 16,274 pieces of evidence across 3,656 anonymous scientific papers, addressing the challenge of grounding comments in long multimodal documents. The dataset is constructed by leveraging explicit references in author rebuttals, providing high‑precision annotations. Evaluation shows that simple retrieval over full paper text performs poorly, evidence‑type inference is a major bottleneck, and multimodal evidence offers complementary signals that pure text misses.

By Serwar Basch, Lizhen Qu, Iryna Gurevych
arXiv Machine Learning
Aug 24

Metag: A dataset to build agentic meta-reviewing capabilities

arXiv:2608.20488v1 Announce Type: new Abstract: AI tools increasingly support tasks across the scientific research cycle, from experiment design and manuscript preparation to peer review. At the same...

By Anirudh Sundar, Min Chen, Divya Tadimeti, Gemma Zhang, Alice Li, Nigel Boachie Kumankumah, Pavan Uttej Ravva, Sadid Hasan, Somya Chatterjee, Pruthvi Prakash Navada, Xiao Wang, Yue Kang, Sulaiman Vesal, Larry Heck