arXiv AI

Inclusion-of-Thoughts: Mitigating Preference Instability via Purifying the Decision Space

arXiv:2604. 04944v2 Announce Type: replace-cross Abstract: Multiple-choice questions (MCQs) are widely used to evaluate large language models (LLMs).

arXiv AI
Sep 3

When Can Large Reasoning Models Save Thinking? Mechanistic Analysis of Behavioral Divergence in Reasoning

The paper investigates why large reasoning models (LRMs) often continue to think even when prompted to stop, a phenomenon called "Still-thinking". By examining confidence at the thinking-termination boundary, internal attention divergences, and attention allocation across prompt segments, the authors find that high perplexity and greater attention to the original question correlate with continued thinking. They propose an attention‑intervention method that suppresses explicit reasoning, which reduces inefficiency but also lowers accuracy, underscoring a trade‑off between instruction compliance, inference speed, and correctness.

By Rongzhi Zhu, Yi Liu, Jiancheng Wang, Xiangyu Liu, Zequn Sun, Yiwei Wang, Yu Deng, Zijian Zhou, Wei Hu
arXiv AI
Sep 16

Can LLMs Model Incorrect Student Reasoning? A Case Study on Distractor Generation

The paper investigates how large language models (LLMs) generate distractor answers for multiple‑choice questions (MCQs) by modeling student misconceptions. It introduces a learning‑science‑based taxonomy of reasoning strategies and applies it to LLM‑generated reasoning traces in math and science MCQs. The study finds that in math, LLMs often follow a misconception‑based process that can be diagnostically useful, whereas in science they rely more on semantic similarity, with frequent failures when the model cannot produce a correct solution or discards plausible distractors. Providing the correct solution in the prompt improves alignment with human distractors by 6.4%. "whyItMatters":"The findings show that anchoring distractor generation to the correct solution enhances LLM alignment with human‑authored distractors, underscoring the importance of correct‑answer cues in educational AI."

By Yanick Zengaffinen, Andreas Opedal, Donya Rooein, Kv Aditya Srivatsa, Shashank Sonkar, Mrinmaya Sachan
arXiv AI
Jul 22

Robust Reasoning Benchmark

arXiv:2604. 08571v3 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) achieve high performance on standard mathematical benchmarks, their problem-solving abilities depend on the context and textual formatting.

By Pavel Golikov, Evgenii Opryshko, Gennady Pekhimenko, Mark C. Jeffrey
arXiv AI
Jun 6

OneReason Technical Report

arXiv:2606. 06260v1 Announce Type: cross Abstract: Generative recommendation models in the OneRec family have been widely deployed in many real-world services, such as short-video, live-streaming, advertising, and e-commerce.

By OneRec Team, Biao Yang, Boyang Ding, Chenglong Chu, Dunju Zang, Fei Pan, Han Li, Hao Jiang, Honghui Bao, Huanjie Wang, Jian Liang, Jiangxia Cao, Jiao Ou, Jiaxin Deng, Jinghao Zhang, Kun Gai, Lu Ren, Peiru Du, Pengfei Zheng, Rongzhou Zhang, Ruiming Tang, Shiyao Wang, Siyang Mao, Siyuan Lou, Teng Shi, Wei Yuan, Wenlong Xu, Xingchen Liu, Xingmei Wang, Xinqi Jin, Yan Sun, Yan Wang, Yifei Hu, Yingzhi He, Yufei Ye, Yuhao Wang, Yunhao Zhou, Yuqin Dai, Zhao Liu, Zhipeng Wei, Zhixin Ling, Ziming Li, Zixing Zhang, Ziyuan Liu, An Zhang, Changxin Lao, Chaoyi Ma, Chengru Song, Defu Lian, Fan Yang, Guowang Zhang, Hao Peng, Jiayao Shen, Jie Chen, Jun Xu, Junmin Chen, Kun Zhang, Kuo Cai, Mingxing Wen, Minmao Wang, Minxuan Lv, Qi Zhang, Qiang Luo, Sheng Yu, Shijie Li, Shijie Yi, Shuang Yang, Shugui Liu, Shuni Chen, Tinghai Zhang, Tingting Gao, Xiang Wang, Xiangyu Wu, Xiangyu Zhao, Xiao Lv, Xiaoyou Zhou, Xuming Wang, Yong Du, Zejian Zhang, Zhaojie Liu, Zhiyang Zhang, Zhuang Zhuang, Ziqi Wang, Ziyi Zhao