arXiv AI

EGTR-Review: Efficient Evidence-Grounded Scientific Peer Review Generation via Multi-Agent Teacher Distillation

arXiv:2606. 06025v1 Announce Type: cross Abstract: Scientific peer review generation has attracted increasing attention for reducing reviewing burdens and providing timely feedback.

arXiv Computation and Language
Sep 3

Expos\'ia: Teaching and Assessment of Academic Writing Skills for Research Project Proposals and Peer Feedback

Exposía is the first public dataset linking academic writing and feedback in higher education, comprising student research project proposals, peer and instructor comments, and free-text reviews collected from a Computer Science course. It includes human assessment scores based on a fine‑grained, pedagogically‑grounded schema for both writing and feedback. The dataset is used to benchmark large language models on automated scoring of proposals and student reviews, revealing that different LLMs excel at each task and that closed‑source models outperform open‑weight ones, while a multi‑aspect prompting strategy proves most effective for classroom deployment.

By Dennis Zyska, Alla Rozovskaya, Ilia Kuznetsov, Iryna Gurevych
arXiv Computation and Language
Sep 16

PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress

PaperDoctor is an agent framework that provides evidence‑grounded, actionable feedback for scientific papers before submission. It evaluates writing, layout, references, code, theory, prior work, and experiments through a three‑layer hierarchical system, linking each critique to specific evidence and revision suggestions. The system selectively rebuilds and reruns experiments to uncover reproducibility gaps, and an interactive interface lets authors explore findings tied to their manuscript.

By Kevin Qinghong Lin, Siyuan Hu, Pan Lu, Yu Chen, Yanzhe Chen, Owen Queen, Yupeng Chen, Jialin Yu, Junchi Yu, Zifeng Ding, Yuanfeng Ji, Sheng Liu, Jindong Gu, Linjie Li, Mike Zheng Shou, Philip Torr, James Zou
arXiv Machine Learning
Sep 21

When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation

Large language models are increasingly used as automated reviewers in scientific evaluation, creating a recursive feedback loop where later reviewers learn from earlier model-generated judgments. A study using Llama 3.1 8B fine‑tuned on ICLR reviews shows that incorporating synthetic reviews compresses rating distributions and reduces semantic diversity, a phenomenon termed scientific‑judgment collapse. To counter this, the authors introduce TrustReviewer, an open‑source LLM system that curates training data and applies paired activation steering at test time to preserve judgment diversity and improve recommendation alignment.

By Sy-Tuyen Ho, Minghui Liu, Furong Huang
arXiv Computation and Language
Sep 3

Reviewing the Reviewer: LLM-Assisted Reviewer Feedback Generation for Guideline Compliance

The paper presents an LLM-driven framework that splits peer reviews into argumentative segments, detects multiple co-occurring issues such as lazy thinking and lack of specificity, and generates targeted, guideline-aware feedback using issue-specific templates. An iterative, reranking-based generation algorithm refines the feedback, and a controlled rewriting study shows it can reduce guideline violations by up to 92.4%. The authors also release LazyReviewPlus, a multi-label dataset of 1,309 sentences annotated for detecting lazy thinking and lack of specificity.

By Sukannya Purkayastha, Qile Wan, Anne Lauscher, Lizhen Qu, Iryna Gurevych
arXiv AI
3d ago

PathAnchor: Path-Structured Evidence for Scientific Agents

PathAnchor is a new scientific reasoning system that uses path-structured evidence workspaces instead of independent passages or concepts. It retrieves source-linked Material‑Sensor‑Signal‑System trajectories that preserve role, direction, and supporting evidence, and a controller uses read‑only tools to search, trace, and open exact evidence before producing a claim‑cited answer. In evaluations on 120 flexible‑sensor questions, PathAnchor achieved an 82.6% score, outperformed six other systems, and improved source recall, citation completeness, and reduced tool calls compared to unordered concept graphs.

By Qiuhui Chen, Jiafan Lu, Shuaimin Tang, Tao Dai, Suyuan Wang, Chenrui Ji, Zhenglei Zhou, Weimin Zhong
arXiv Computation and Language
Sep 22

To Consolidate or not to Consolidate? Evaluating the Impact of Consolidation in Multi-Reference Training using Peer Reviews

arXiv:2609.22805v1 Announce Type: new Abstract: Natural language generation (NLG) tasks span the spectrum of conditional entropy, ranging from highly constrained machine translation to open-ended dia...

By Maitreya Prafulla Chitale, Ketaki Mangesh Shetye, Yash More, Harshit Gupta, Manav Chaudhary, Manish Shrivastava, Vasudeva Varma