arXiv:2602. 05395v2 Announce Type: replace-cross Abstract: A simple strategy for improving LLM accuracy, especially in math and reasoning problems, is to sample multiple responses and submit the answer most consistently reached.
By Jingkai Huang, Will Ma, Zhengyuan Zhou
Using LLMs as judges has become standard practice for evaluating model outputs at scale. This is particularly common for subjective, open-ended tasks such as assessing helpfulness or alignment, where no single reference answer exists.
arXiv:2602. 13935v2 Announce Type: replace Abstract: While LLMs have seen substantial improvement in reasoning capabilities, they also sometimes overthink, generating unnecessary reasoning steps, particularly under uncertainty, given ill-posed or ambiguous queries.
By Yangxinyu Xie, Tao Wang, Soham Mallick, Yan Sun, Georgy Noarov, Mengxin Yu, Tanwi Mallick, Weijie J. Su, Edgar Dobriban
arXiv:2606. 20820v2 Announce Type: replace Abstract: Can we trust evaluation scores to capture an LLM's true real-world performance?
By Zhijian Zhou, Zesheng Ye, Zhaorun Chen, Bo Li, Feng Liu
arXiv:2607. 27023v1 Announce Type: new Abstract: Evaluating large generative models across benchmarks is time-consuming and computationally expensive.
By Paula Cordero Encinar, Taylan Cemgil, Arnaud Doucet, Virginia Aglietti, Silvia Chiappa
arXiv:2607. 03991v1 Announce Type: new Abstract: Repeated LLM calls are the standard way to estimate how trustworthy a Text-to-SQL result is: run the pipeline multiple times, judge each SQL execution, and use the consistency of the verdicts as a confidence signal.
By Yaron Anavi, Mor Aisenberg, Nadav Nesher, Elena Khabibullina, Isabella Cattinelli