arXiv AI

Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable

arXiv:2603. 20450v2 Announce Type: replace-cross Abstract: A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polishing, paraphrasing, and grammar correction of otherwise human-written reviews.

Hugging Face Trending Papers
Aug 4

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are. We address these questions by first surveying reviewer-facing AI policies across 111 leading AI/NLP conferences and medical journals, revealing substantial regulation differences between the two communities.

arXiv AI
Aug 5

AI-Assisted Peer Review Across Research Communities: From Reviewer AI Policies to LLM Review Quality

arXiv:2608. 03581v1 Announce Type: cross Abstract: AI-assisted peer review is increasingly discussed and adopted as a tool to support the scientific publishing process, yet there is little systematic understanding of how publication venues regulate its use or of how capable current AI review systems are.

By Alexander M. Fichtl, Lukas Ellinger, Josefin Kelber, Kry\v{s}tof Ol\'ik, Georg Groh
arXiv AI
3d ago

Policy-Conditioned AI-Use Detection: An Evidentiary Framework for Academic Publishing

The paper introduces a policy‑conditioned AI‑use detection framework for academic publishing, arguing that traditional AI detection tools misalign with venue rules by merely identifying AI‑generated text. Instead, the proposed system treats the governing policy as an explicit input, generating hypotheses, evidence, calibration, and uncertainty rather than binary verdicts. It outlines how to benchmark compliance, evaluate true positive rates at venue‑specified false positive thresholds, and emphasizes the need for structured disclosure, tool routing, and contestable findings.

By Jairo Diaz-Rodriguez, Mumin Jia
arXiv AI
Aug 19

Hidden Prompts in Manuscripts Exploit AI-Assisted Peer Review

The article reports that in July 2025, 18 arXiv manuscripts contained hidden instructions designed to manipulate AI‑assisted peer review, such as covert commands to give only positive reviews. These prompts were concealed using white text and microscopic fonts, and the authors’ reactions ranged from withdrawal to defending the practice as a test of reviewer misuse of large language models. The study identifies four types of hidden prompts, critiques the ineffectiveness of honeypot defenses, and highlights inconsistent publisher policies while calling for controlled AI integration and harmonized guidelines in academic evaluation.

By Zhicheng Lin
arXiv Computation and Language
Sep 24

How Much Were You Told? Measuring External Information in Peer Reviews

The paper introduces Self‑Conditioning, an unsupervised, information‑theoretic estimator that measures the amount of external information in peer reviews. It compares the likelihood of a review under its original production context with the likelihood when that context is augmented by hints extracted from the review itself. On the IntelLabs benchmark, Self‑Conditioning can perfectly distinguish fully‑delegated reviews from machine‑polished ones, remains largely insensitive to surface rewriting, and shows that increased external input drives scores toward human‑like values, unlike standard ATD baselines.

By Matthieu Dubois, Pablo Piantanida, Fran\c{c}ois Yvon
arXiv AI
Sep 4

HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews

HalluPeer is a new benchmark designed to detect hallucinations in scientific peer reviews. It provides aligned triples of paper content, human-written reviews, and hallucination-injected reviews, annotated for detection, classification, and localization. Experiments on 12K papers and 38K reviews show that existing detectors struggle to separate hallucinations from legitimate critique, and that HalluPeer-defined hallucination patterns occur in real peer reviews.

By Tzu-Ling Lin, Dong-Ting Yao, Teng-Fang Hsiao, Wei-Chih Chen, Hong-Han Shuai
Hugging Face Trending Papers
Sep 3

HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews

HalluPeer is a new benchmark designed to detect hallucinations in scientific peer reviews. It provides aligned triples of paper content, human-written reviews, and hallucination-injected reviews, annotated for detection, classification, and localization. Experiments on 12K papers and 38K reviews show that current detectors struggle to distinguish hallucinations from legitimate critique, and real peer reviews contain HalluPeer-defined hallucination patterns, underscoring the need for source-aware verification.

arXiv AI
Jun 17

BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?

arXiv:2510. 18003v2 Announce Type: replace-cross Abstract: The convergence of LLM-powered research assistants and AI-based peer review systems creates a critical vulnerability: fully automated publication loops where AI-generated research is evaluated by AI reviewers without human oversight.

By Fengqing Jiang, Yichen Feng, Yuetai Li, Luyao Niu, Basel Alomair, Radha Poovendran