When people share experiences online, they often express thoughts in two ways: a star rating and a written review. In sentiment analysis, ratings are widely used as convenient weak labels for textual sentiment, yet whether the two actually agree is rarely questioned.
arXiv:2606. 16344v1 Announce Type: new Abstract: Travelers increasingly ask large language model (LLM) assistants which hotel to book, making these systems gatekeepers of property visibility -- yet what moves their recommendations is undocumented.
By Mirza Samad Ahmed Baig, Syeda Anshrah Gillani, Asher Ali
arXiv:2606. 15887v1 Announce Type: cross Abstract: Large language model (LLM) systems are increasingly proposed to assist peer review, yet most evaluations judge the prose of machine-generated review text, not the validity of the numeric score a system assigns.
By Costa Georgantas
arXiv:2607. 14174v1 Announce Type: new Abstract: Financial sentiment extraction has largely relied on news text and supervised extraction against return labels alone, leaving 10-K filings -- and volatility, the target risk disclosure is arguably best suited to informing -- comparatively unexplored.
By Sanggyu Sean Choi
arXiv:2608. 08975v1 Announce Type: cross Abstract: As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions.
By Ming Li, Chenguang Wang, Xirui Li, Xinyue Zeng, Dianqi Li, Peng Shi, Dawei Zhou, Tianyi Zhou
arXiv:2205. 11930v3 Announce Type: replace-cross Abstract: Human ratings are the gold standard in NLG evaluation.
By Kawin Ethayarajh, Dan Jurafsky