arXiv:2510. 08647v2 Announce Type: replace-cross Abstract: Recent developments have enabled advanced reasoning in Large Language Models (LLMs) via long Chain-of-Thought (CoT), trading efficiency during inference for performance.
By Chengzhengxu Li, Xiaoming Liu, Zhaohan Zhang, Shengchao Liu, Guoxin Ma, Yu Lan, Cong Wang, Chao Shen
arXiv:2608. 03291v1 Announce Type: cross Abstract: Chain-of-thought (CoT) reasoning improves large language model (LLM) performance while also providing an observable interface to the model's reasoning process.
By Shashwat Sourav, Aishwarya Balwani
The paper introduces the Flow Moment, a reasoning pattern marked by sustained, process‑confirming verbalizations, contrasting with the revision‑oriented Aha Moment. It proposes Flow‑CoT, a rewritten version of reasoning traces that preserves content while highlighting Flow Markers, and uses it as auxiliary supervision in on‑policy self‑distillation (OPSD). The authors further present Aha‑Flow Distillation (AFD), a dual‑mode extension of OPSD that pairs concise solution‑based supervision (Aha branch) with rewritten Flow‑CoT under a confident reasoning instruction (Flow branch). Experiments on AIME25 and HMMT25 with Qwen3‑8B and Qwen3‑4B models show consistent performance gains, and controlled ablations confirm that the dual‑mode training structure contributes to the improvement.
By Xiaodong Wang, Peixi Peng
The study investigates how deictic ambiguity—specifically the shifting reference of expressions like "previous"—affects Draft‑Verify‑Revise pipelines that use multiple large language models (LLMs). Using a synthetic dataset of 10 base examples and 21 reasoning‑effort configurations, six LLMs were evaluated for their ability to correctly resolve the ambiguous expression across the draft, verify, and revise stages. Results show wide variance in balanced accuracy, with GPT‑5.2 improving from 0.156 to 0.942 with increased reasoning effort, while Gemini 3 Pro consistently achieved high accuracy above 0.94 even at low reasoning effort, and meta‑evaluators often relied on surface cues when making errors.
By Obinna I. Ekekezie
arXiv:2607. 22629v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) produce long, explicit chains of intermediate steps before generating a final answer at inference time.
By Durgesh Kalwar, Vardhan Palod, Subbarao Kambhampati
arXiv:2606. 05402v1 Announce Type: cross Abstract: Large reasoning models (LRMs) produce reasoning traces with non-linear structures, such as backtracking and self-correction, that complicate the evaluation and monitoring of the reasoning process.
By Jinu Lee, Shivam Agarwal, Amruta Parulekar, Siddarth Madala, Dilek Hakkani-Tur, Julia Hockenmaier