arXiv:2604. 11996v2 Announce Type: replace-cross Abstract: Should we trust Large Language Models (LLMs) with high accuracy?
By Manas Pathak, Xingyao Chen, Shuozhe Li, Amy Zhang, Liu Leqi
arXiv:2603. 05167v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used as judges of chain-of-thought (CoT) reasoning, yet it remains unclear whether they can reliably assess process faithfulness rather than merely answer plausibility.
By Avni Mittal, Rauno Arike
arXiv:2606. 31543v1 Announce Type: new Abstract: Large language models can produce fluent, internally coherent reasoning traces for abstract reasoning tasks while still being confidently wrong - making selection among candidates, not just generation, the central challenge.
By Johan Land
arXiv:2607. 10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution.
By Ning Liu
arXiv:2608.29168v1 Announce Type: new
Abstract: The LLM-as-a-Judge paradigm has emerged as a scalable alternative to human evaluation. However, single-model judges are limited by their inherent model...
By Yiyue Qian, Shinan Zhang, Huan Song, Hannah Marlowe
arXiv:2609.38409v1 Announce Type: new
Abstract: Recent progress in large language model reasoning has been driven by benchmarks and reinforcement learning environments with automatically verifiable r...
By \.Ibrahim Ethem Deveci, Funda Tan \c{C}al{\i}k, Bar{\i}\c{s} Deniz Sa\u{g}lam, Duygu Ataman