arXiv:2608. 08975v1 Announce Type: cross Abstract: As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions.
By Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou
arXiv:2609.39027v1 Announce Type: new
Abstract: AI reviewers can assign different judgments to manuscripts that report the same science in different wording, potentially rewarding rhetorical optimiza...
By Chenguang Wang, Ming Li, Chengrui Fan, Jianpeng Chen, Han Chen, Tianyi Zhou, Dawei Zhou
arXiv:2609.23264v1 Announce Type: new
Abstract: Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high sc...
By Shakiba Amirshahi, Sajad Ebrahimi, Hai Son Le, Negar Arabzadeh, Ebrahim Bagheri
Large language models (LLMs) are widely used to assist writing, but this study shows they alter both tone and meaning of human text. A user study found that heavy LLM use increased neutral essays by nearly 70% and made writers feel less creative and less in their voice. Even when prompted to make only grammar edits, LLMs changed the semantic content of essays and produced AI-generated scientific reviews that were less focused on clarity and significance and scored higher on average.
By Marwa Abdulhai, Isadora White, Yanming Wan, Ibrahim Qureshi, Joel Z. Leibo, Max Kleiman-Weiner, Natasha Jaques
arXiv:2606. 10159v1 Announce Type: cross Abstract: AI is increasingly used to support scientific peer review, from manuscript screening, reviewer assistance to editorial triage.
By Lin Li, Qi Zhang, Xander Davies, Jianing Qiu, Yarin Gal
arXiv:2607. 21498v1 Announce Type: cross Abstract: A rhetorical figure that Cicero and Quintilian catalogued two thousand years ago reappears, systematically, in the text of large language models: epanorthosis, the self-correction of the specimen {\guillemotleft}This is not a course.
By Federico Boggia
A rhetorical figure that Cicero and Quintilian catalogued two thousand years ago reappears, systematically, in the text of large language models: epanorthosis, the self-correction of the specimen «This is not a course. It is a journey of transformation».
arXiv:2608. 14630v1 Announce Type: cross Abstract: Human decision-making is often shaped by a range of well-documented cognitive biases.
By Zirui Cheng, Joey Chan, Simo Du, Chenhao Tan, Yue Guo, Hao Peng
arXiv:2603. 20450v2 Announce Type: replace-cross Abstract: A number of scientific conferences and journals have recently enacted policies that prohibit LLM usage by peer reviewers, except for polishing, paraphrasing, and grammar correction of otherwise human-written reviews.
By Rounak Saha, Gurusha Juneja, Dayita Chaudhuri, Naveeja Sajeevan, Nihar B Shah, Danish Pruthi
The study investigates bias in large language model (LLM) judges by having ten LLMs evaluate narrative constraint selections rather than generated text. Results show that self-preference largely disappears under blind evaluation when quality and evaluator severity are controlled, but self- and other-labels alone shift scores bidirectionally when quality is matched. The authors conclude that authorship attribution drives evaluation bias and that open-ended, ground‑truth‑free tasks can effectively study LLM judge behavior.
By Songeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung
Large language models are increasingly used as automated reviewers in scientific evaluation, creating a recursive feedback loop where later reviewers learn from earlier model-generated judgments. A study using Llama 3.1 8B fine‑tuned on ICLR reviews shows that incorporating synthetic reviews compresses rating distributions and reduces semantic diversity, a phenomenon termed scientific‑judgment collapse. To counter this, the authors introduce TrustReviewer, an open‑source LLM system that curates training data and applies paired activation steering at test time to preserve judgment diversity and improve recommendation alignment.
By Sy-Tuyen Ho, Minghui Liu, Furong Huang
arXiv:2609.08016v1 Announce Type: new
Abstract: Multi-agent debate, in which several LLMs exchange arguments before answering, is widely assumed to improve answer quality by surfacing genuine disagre...
By Chen Qian