arXiv AI By Jiaqi Xu, Cuiling Lan, Xuejin Chen, Yan Lu

Stepwise Think-Critique: Interleaved Reasoning and Self-Critique in a Single LLM

Read the original on arXiv AI →

The paper introduces Stepwise Think-Critique (STC), an end‑to‑end trainable framework that lets a single large language model (LLM) interleave reasoning with inline, step‑level critique. STC uses reinforcement learning to jointly reward reasoning accuracy and critique consistency, achieving a 7.2% improvement in Pass@1 on five mathematical reasoning benchmarks and a 67.4% step‑level critique F1 score, outperforming several external reward models. This approach aligns LLM behavior more closely with human critical thinking by embedding self‑evaluation directly into the reasoning process.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 4

Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning

The paper investigates whether the text of chain‑of‑thought reasoning steps actually reflects their true importance for a model’s final answer. By defining step importance as the advantage in expected reward when a step is included, the authors use Monte Carlo rollouts to estimate ground truth and then test whether large language model judges can identify high‑advantage steps. They find that capable LLMs can beat a prevalence baseline but still fall far short of a noise ceiling, and that fine‑tuning a step‑level critic improves detection for incorrect responses but remains distant from the ceiling for correct ones, indicating that step importance is only partially recoverable from the reasoning trace text.

By Kevin Du, Alexander Hoyle, Laura Ruis, Acyr Locatelli
arXiv AI
Jul 28

Offline-Online Curriculum RL for Multimodal Reasoning

arXiv:2607. 23700v1 Announce Type: new Abstract: Multimodal large language models exhibit capabilities on reasoning tasks, yet often produce flawed intermediate steps while yielding correct final answers.

By Wendi Deng, Hang Du, Guoshun Nan, Haokun Tian, Jiaqi Yu, Xinlei Cao, Jaile Li, Jingfeng Chen, Ling Deng, Ting Li, Hao Yang, Jun Liu, Xudong Jiang, Sicong Leng