arXiv:2609.08475v1 Announce Type: cross
Abstract: Large language models have collapsed the cost of producing lexically elaborate prose, and whether peer reviewers still reward it is a question about...
By Jiabin Zheng (School of Computer Science, Peking University)
The paper introduces Self‑Conditioning, an unsupervised, information‑theoretic estimator that measures the amount of external information in peer reviews. It compares the likelihood of a review under its original production context with the likelihood when that context is augmented by hints extracted from the review itself. On the IntelLabs benchmark, Self‑Conditioning can perfectly distinguish fully‑delegated reviews from machine‑polished ones, remains largely insensitive to surface rewriting, and shows that increased external input drives scores toward human‑like values, unlike standard ATD baselines.
By Matthieu Dubois, Pablo Piantanida, Fran\c{c}ois Yvon
When people share experiences online, they often express thoughts in two ways: a star rating and a written review. In sentiment analysis, ratings are widely used as convenient weak labels for textual sentiment, yet whether the two actually agree is rarely questioned.
arXiv:2606. 16344v1 Announce Type: new Abstract: Travelers increasingly ask large language model (LLM) assistants which hotel to book, making these systems gatekeepers of property visibility -- yet what moves their recommendations is undocumented.
By Mirza Samad Ahmed Baig, Syeda Anshrah Gillani, Asher Ali
arXiv:2609.23264v1 Announce Type: new
Abstract: Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high sc...
By Shakiba Amirshahi, Sajad Ebrahimi, Hai Son Le, Negar Arabzadeh, Ebrahim Bagheri
arXiv:2606. 15887v1 Announce Type: cross Abstract: Large language model (LLM) systems are increasingly proposed to assist peer review, yet most evaluations judge the prose of machine-generated review text, not the validity of the numeric score a system assigns.
By Costa Georgantas
arXiv:2607. 14174v1 Announce Type: new Abstract: Financial sentiment extraction has largely relied on news text and supervised extraction against return labels alone, leaving 10-K filings -- and volatility, the target risk disclosure is arguably best suited to informing -- comparatively unexplored.
By Sanggyu Sean Choi
arXiv:2608. 08975v1 Announce Type: cross Abstract: As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions.
By Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou
The paper critiques the conventional method of averaging human ratings to evaluate natural language generation (NLG) systems, arguing that it relies on assumptions about annotators that are often violated, especially when using Likert scales. These violations can even reverse true preferences, leading to inaccurate system rankings. The authors propose a more theoretically sound protocol and introduce a new system-level probabilistic assessment (SPA) for open-ended tasks like story generation, which successfully recovers expected model orderings where the standard protocol fails.
By Kawin Ethayarajh, Dan Jurafsky
The study investigates how prior scores influence large language model (LLM) judgments in the LLM-as-a-Judge paradigm. By testing three prompt conditions—no metadata, revision framing, and anchored metadata containing prior scores—the authors find that prior scores systematically bias evaluations, shifting ratings toward those scores across 192,000 attempts. The bias also affects categorical decisions, blocking 48% of error corrections and flipping 10.18% of correct judgments, and is not mitigated by Chain-of-Thought or a warning, underscoring the need for careful context engineering.
By Ante Kapetanovic, Kemal Altwlkany, Andro Mercep, Tomislav Duricic, Emanuel Lacic
arXiv:2607. 29516v1 Announce Type: cross Abstract: AI coding agents are generating code at volumes that exceed the capacity of traditional peer review.
By Chandra Maddila, Mashrur Rashik, Euna Mehnaz Khan, Smriti Jha, James Saindon, Nachi Nagappan, Peter C. Rigby
arXiv:2608. 10008v1 Announce Type: cross Abstract: LLM recommenders for top-$K$ item suggestion regularly emit titles outside the target catalog.
By Srijith Ravikumar