arXiv:2608.29257v1 Announce Type: new
Abstract: Multiple-choice questions (MCQs) are a standard format for evaluating large language models (LLMs), yet the popularity of answer options can confound e...
By Abdelrahman Abdallah, Mohammed Ali, Bhawna Piryani, Mahmoud Abdalla, Adam Jatowt
arXiv:2607.14109v2 Announce Type: replace
Abstract: Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central...
By Inder Preet, Shuxin Lin, Dhaval Patel
arXiv:2607. 14109v1 Announce Type: cross Abstract: Probing the capabilities of Large Language Models (LLMs) and building robust solutions for Multiple-Choice Question Answering (MCQA) remain central challenges in natural language understanding.
By Inder Preet, Shuxin Lin, Dhaval Patel
The paper investigates why large reasoning models (LRMs) often continue to think even when prompted to stop, a phenomenon called "Still-thinking". By examining confidence at the thinking-termination boundary, internal attention divergences, and attention allocation across prompt segments, the authors find that high perplexity and greater attention to the original question correlate with continued thinking. They propose an attention‑intervention method that suppresses explicit reasoning, which reduces inefficiency but also lowers accuracy, underscoring a trade‑off between instruction compliance, inference speed, and correctness.
By Rongzhi Zhu, Yi Liu, Jiancheng Wang, Xiangyu Liu, Zequn Sun, Yiwei Wang, Yu Deng, Zijian Zhou, Wei Hu
arXiv:2608. 15065v1 Announce Type: new Abstract: Large Reasoning Models produce diverse, sometimes inconsistent answers across repeated queries on the same problem, so multi-sample inference is a prerequisite for reliable deployment.
By Chanhee Park, Sungbin Han, Jeongho Yoon, Seongtae Hong, Heuiseok Lim
The paper investigates how large language models (LLMs) generate distractor answers for multiple‑choice questions (MCQs) by modeling student misconceptions. It introduces a learning‑science‑based taxonomy of reasoning strategies and applies it to LLM‑generated reasoning traces in math and science MCQs. The study finds that in math, LLMs often follow a misconception‑based process that can be diagnostically useful, whereas in science they rely more on semantic similarity, with frequent failures when the model cannot produce a correct solution or discards plausible distractors. Providing the correct solution in the prompt improves alignment with human distractors by 6.4%.
"whyItMatters":"The findings show that anchoring distractor generation to the correct solution enhances LLM alignment with human‑authored distractors, underscoring the importance of correct‑answer cues in educational AI."
By Yanick Zengaffinen, Andreas Opedal, Donya Rooein, Kv Aditya Srivatsa, Shashank Sonkar, Mrinmaya Sachan
arXiv:2604. 08571v3 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) achieve high performance on standard mathematical benchmarks, their problem-solving abilities depend on the context and textual formatting.
By Pavel Golikov, Evgenii Opryshko, Gennady Pekhimenko, Mark C. Jeffrey
arXiv:2604. 01161v2 Announce Type: replace Abstract: Large language models (LLMs) exhibiting test-time scaling behavior, such as extended reasoning traces and self-verification, have demonstrated remarkable performance on complex, long-term reasoning tasks.
By Gleb Rodionov, Roman Garipov, George Yakushev
arXiv:2606. 06260v1 Announce Type: cross Abstract: Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce.
By OneRec Team, Biao Yang, Boyang Ding, Chenglong Chu, Dunju Zang, Fei Pan, Han Li, Hao Jiang, Honghui Bao, Huanjie Wang, Jian Liang, Jiangxia Cao, Jiao Ou, Jiaxin Deng, Jinghao Zhang, Kun Gai, Lu Ren, Peiru Du, Pengfei Zheng, Rongzhou Zhang, Ruiming Tang, Shiyao Wang, Siyang Mao, Siyuan Lou, Teng Shi, Wei Yuan, Wenlong Xu, Xingchen Liu, Xingmei Wang, Xinqi Jin, Yan Sun, Yan Wang, Yifei Hu, Yingzhi He, Yufei Ye, Yuhao Wang, Yunhao Zhou, Yuqin Dai, Zhao Liu, Zhipeng Wei, Zhixin Ling, Ziming Li, Zixing Zhang, Ziyuan Liu, An Zhang, Changxin Lao, Chaoyi Ma, Chengru Song, Defu Lian, Fan Yang, Guowang Zhang, Hao Peng, Jiayao Shen, Jie Chen, Jun Xu, Junmin Chen, Kun Zhang, Kuo Cai, Mingxing Wen, Minmao Wang, Minxuan Lv, Qi Zhang, Qiang Luo, Sheng Yu, Shijie Li, Shijie Yi, Shuang Yang, Shugui Liu, Shuni Chen, Tinghai Zhang, Tingting Gao, Xiang Wang, Xiangyu Wu, Xiangyu Zhao, Xiao Lv, Xiaoyou Zhou, Xuming Wang, Yong Du, Zejian Zhang, Zhaojie Liu, Zhiyang Zhang, Zhuang Zhuang, Ziqi Wang, Ziyi Zhao
arXiv:2607. 18100v1 Announce Type: new Abstract: Extended reasoning has become standard for frontier Large Language Models (LLMs), yet the trajectories these models produce remain largely uncontrollable.
By Sheldon Yu, Tong Yu, Xunyi Jiang, Rohan Surana, Gagan Mundada, Sungchul Kim, Lina Yao, Julian McAuley, Junda Wu
arXiv:2606. 16360v1 Announce Type: cross Abstract: Chain-of-thought (CoT) prompting improves reasoning in large language models (LLMs) by externalizing intermediate computation as discrete text tokens, but this textual interface also introduces redundancy and inference overhead.
By Hanyu Lin, Min Cai, Jiawei Wen, Haodi Zhang
arXiv:2605. 26795v2 Announce Type: replace Abstract: Chain-of-thought (CoT) prompting enhances large language model performance, yet what drives these gains remains unclear.
By Xiang Wang, Wei Wei