arXiv:2604. 11996v2 Announce Type: replace-cross Abstract: Should we trust Large Language Models (LLMs) with high accuracy?
By Manas Pathak, Xingyao Chen, Shuozhe Li, Amy Zhang, Liu Leqi
arXiv:2606. 31543v1 Announce Type: new Abstract: Large language models can produce fluent, internally coherent reasoning traces for abstract reasoning tasks while still being confidently wrong - making selection among candidates, not just generation, the central challenge.
By Johan Land
arXiv:2606. 05402v1 Announce Type: cross Abstract: Large reasoning models (LRMs) produce reasoning traces with non-linear structures, such as backtracking and self-correction, that complicate the evaluation and monitoring of the reasoning process.
By Jinu Lee, Shivam Agarwal, Amruta Parulekar, Siddarth Madala, Dilek Hakkani-Tur, Julia Hockenmaier
The paper introduces LLM‑PeerReview, an unsupervised ensemble method that selects the best response from multiple LLM-generated candidates by scoring each answer with several LLMs, aggregating those scores via averaging or a graphical model, and choosing the highest-scoring response. The approach is peer‑review inspired, transparent, and interpretable, and it outperforms the Smoothie‑Global model by 6.9%–7.3% across factual recall QA, math reasoning, and instruction‑following tasks. The authors also provide a curated benchmark suite of 12 ensemble methods evaluated on four datasets and three task families to aid reproducibility.
By Zhijun Chen, Zeyu Ji, Qianren Mao, Hao Wu, Jinhuan Song, Junhang Cheng, Bangjie Qin, Zhuoran Li, Jingzheng Li, Kai Sun, Zizhe Wang, Yikun Ban, Zhu Sun, Xiangyang Ji, Hailong Sun, Xiao Huang
arXiv:2604.14121v3 Announce Type: replace
Abstract: Large language models (LLMs) have become increasingly used for various tasks, often coupled with Chain-of-Thought (CoT) prompting to boost accuracy...
By Zipeng Ling, Shuliang Liu, Seonil Son, Shenghong Fu, Yuehao Tang, Yao Wan, Xuming Hu
arXiv:2608. 15303v1 Announce Type: new Abstract: Test-time compute can substantially improve Large Language Model (LLM) reasoning performance, yet how and when additional compute helps remains poorly understood.
By Bo Wen, Yuhao Chen, Erhan Bilal, Carla Agurto Rios, Chen Wang, Junchen Jiang