arXiv:2609.23264v1 Announce Type: new
Abstract: Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high sc...
By Shakiba Amirshahi, Sajad Ebrahimi, Hai Son Le, Negar Arabzadeh, Ebrahim Bagheri
arXiv:2609.05947v1 Announce Type: new
Abstract: Peer review is central to quality control in science. However, existing evaluations of AI-assisted peer review mainly focus on the overall quality of g...
By Siming Yuan, Xueyi Zhang, Wangze Ni, Tianfang Xiao, Shimin Di, Jia Zhu, Zhuoren Jiang, Rong Tan, Lei Chen, Kui Ren
arXiv:2608.28626v1 Announce Type: cross
Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This st...
By Emad Alharbi
arXiv:2609.08475v1 Announce Type: cross
Abstract: Large language models have collapsed the cost of producing lexically elaborate prose, and whether peer reviewers still reward it is a question about...
By Jiabin Zheng (School of Computer Science, Peking University)
The paper argues that verbalized confidence—once viewed as overconfident and coarse—has become the preferred soft‑scoring method for LLM‑as‑a‑Judge on top‑tier proprietary models released after 2025. Experiments on SummEval, AggreFact, and HelpSteer2 across up to 18 LLMs show that log‑probabilities are no longer the best signal, and that adding an overconfidence advisory and self‑debate further improves calibration and robustness. The authors note that these enhancements incur little accuracy loss on post‑2025 models but do affect pre‑2025 ones, highlighting a compatibility shift in how confidence should be measured.
By Yu-Chung Hsiao
arXiv:2609.13824v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faste...
By Aakash Kumar Tiwari
arXiv:2609.13773v1 Announce Type: new
Abstract: LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may...
By Ruichen Zheng, Yihe Wang, Fabrice Y Harel-Canada, Sara Khosravi, Zeynep Senahan Yildiz, Amit Sahai, Nanyun Peng
arXiv:2610.01471v1 Announce Type: cross
Abstract: Large language models now generate code, documentation, and analyses, and are increasingly used to review such output. We ask when a second review by...
By Tae-Eun Song
arXiv:2608.21374v1 Announce Type: new
Abstract: Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspec...
By Ruotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue, Dong Liang, Jigao Fu, Wu YanBiao, Yuanyi Zhen, Fengli Xu, Yong Li
arXiv:2606. 07810v1 Announce Type: cross Abstract: Large language models (LLMs) are widely used as judges for evaluating model outputs, but their high cost, latency, and opacity limit scalability.
By Anish Laddha, Nitesh Pradhan, Gaurav Srivastava
The study investigates how prior scores influence large language model (LLM) judgments in the LLM-as-a-Judge paradigm. By testing three prompt conditions—no metadata, revision framing, and anchored metadata containing prior scores—the authors find that prior scores systematically bias evaluations, shifting ratings toward those scores across 192,000 attempts. The bias also affects categorical decisions, blocking 48% of error corrections and flipping 10.18% of correct judgments, and is not mitigated by Chain-of-Thought or a warning, underscoring the need for careful context engineering.
By Ante Kapetanovic, Kemal Altwlkany, Andro Mercep, Tomislav Duricic, Emanuel Lacic
arXiv:2608. 08975v1 Announce Type: cross Abstract: As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions.
By Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou