Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation
arXiv:2606. 07936v1 Announce Type: cross Abstract: Human evaluation plays a critical role in assessing the quality of generated text.
arXiv:2510. 02027v2 Announce Type: replace Abstract: Scholarly publishing requires scalable scrutiny supported by auditable evidence.
arXiv:2606. 07936v1 Announce Type: cross Abstract: Human evaluation plays a critical role in assessing the quality of generated text.
arXiv:2608. 03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting.
arXiv:2607. 16989v1 Announce Type: cross Abstract: Introduction.
arXiv:2606. 06025v1 Announce Type: cross Abstract: Scientific peer review generation has attracted increasing attention for reducing reviewing burdens and providing timely feedback.
arXiv:2507. 06185v2 Announce Type: replace-cross Abstract: In July 2025, 18 academic manuscripts on arXiv contained hidden instructions that manipulated AI-assisted peer review (indirect prompt injection).
arXiv:2608. 14625v1 Announce Type: cross Abstract: Academic peer review is under mounting strain: NeurIPS 2025 received 21,575 submissions, ICLR 2025 received 11,603, and ICML 2025 received 12,107.
arXiv:2604. 04074v4 Announce Type: replace Abstract: Large language model (LLM)-based reviewing systems typically assess manuscripts in isolation, leaving literature- and code-dependent claims difficult to verify.
AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are. We address these questions by first surveying reviewer-facing AI policies across 111 leading AI/NLP conferences and medical journals, revealing substantial regulation differences between the two communities.
arXiv:2606. 02255v1 Announce Type: cross Abstract: Human annotation is the empirical foundation of much NLP research, from dataset construction to model evaluation, but papers often leave unclear who produced the annotations and how the annotation process was controlled.
arXiv:2608. 03581v1 Announce Type: cross Abstract: AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are.
Generative-AI evaluations can become historical before publication, yet calendar age does not affect every conclusion equally. This paper has two linked purposes.
arXiv:2607. 26066v1 Announce Type: cross Abstract: The growing volume of scientific submissions has motivated interest in using large language models (LLMs) to assist peer review.