arXiv AI

Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction

arXiv:2607. 11696v1 Announce Type: new Abstract: Self-refinement often fails to strengthen few-shot inductive reasoning in large language models.

arXiv AI
Sep 10

Aha-Flow Distillation: Flow Markers Matter in LLM Reasoning

The paper introduces the Flow Moment, a reasoning pattern marked by sustained, process‑confirming verbalizations, contrasting with the revision‑oriented Aha Moment. It proposes Flow‑CoT, a rewritten version of reasoning traces that preserves content while highlighting Flow Markers, and uses it as auxiliary supervision in on‑policy self‑distillation (OPSD). The authors further present Aha‑Flow Distillation (AFD), a dual‑mode extension of OPSD that pairs concise solution‑based supervision (Aha branch) with rewritten Flow‑CoT under a confident reasoning instruction (Flow branch). Experiments on AIME25 and HMMT25 with Qwen3‑8B and Qwen3‑4B models show consistent performance gains, and controlled ablations confirm that the dual‑mode training structure contributes to the improvement.

By Xiaodong Wang, Peixi Peng
arXiv Computation and Language
Sep 14

Can LLMs in Draft-Verify-Revise Pipelines Resolve Deictic Ambiguity?

The study investigates how deictic ambiguity—specifically the shifting reference of expressions like "previous"—affects Draft‑Verify‑Revise pipelines that use multiple large language models (LLMs). Using a synthetic dataset of 10 base examples and 21 reasoning‑effort configurations, six LLMs were evaluated for their ability to correctly resolve the ambiguous expression across the draft, verify, and revise stages. Results show wide variance in balanced accuracy, with GPT‑5.2 improving from 0.156 to 0.942 with increased reasoning effort, while Gemini 3 Pro consistently achieved high accuracy above 0.94 even at low reasoning effort, and meta‑evaluators often relied on surface cues when making errors.

By Obinna I. Ekekezie