arXiv:2609.23264v1 Announce Type: new
Abstract: Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high sc...
By Shakiba Amirshahi, Sajad Ebrahimi, Hai Son Le, Negar Arabzadeh, Ebrahim Bagheri
arXiv:2609.05947v1 Announce Type: new
Abstract: Peer review is central to quality control in science. However, existing evaluations of AI-assisted peer review mainly focus on the overall quality of g...
By Siming Yuan, Xueyi Zhang, Wangze Ni, Tianfang Xiao, Shimin Di, Jia Zhu, Zhuoren Jiang, Rong Tan, Lei Chen, Kui Ren
arXiv:2608.28626v1 Announce Type: cross
Abstract: Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This st...
By Emad Alharbi
arXiv:2609.08475v1 Announce Type: cross
Abstract: Large language models have collapsed the cost of producing lexically elaborate prose, and whether peer reviewers still reward it is a question about...
By Jiabin Zheng (School of Computer Science, Peking University)
The paper argues that verbalized confidence—once viewed as overconfident and coarse—has become the preferred soft‑scoring method for LLM‑as‑a‑Judge on top‑tier proprietary models released after 2025. Experiments on SummEval, AggreFact, and HelpSteer2 across up to 18 LLMs show that log‑probabilities are no longer the best signal, and that adding an overconfidence advisory and self‑debate further improves calibration and robustness. The authors note that these enhancements incur little accuracy loss on post‑2025 models but do affect pre‑2025 ones, highlighting a compatibility shift in how confidence should be measured.
By Yu-Chung Hsiao
arXiv:2609.13824v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faste...
By Aakash Kumar Tiwari