arXiv:2508. 09883v2 Announce Type: replace-cross Abstract: Large language models (LLMs) demonstrate remarkable reasoning capabilities in tasks such as algorithmic coding and mathematical problem-solving.
By Xiaojun Wu, Xiaoguang Jiang, Huiyang Li, Jucai Zhai, Dengfeng Liu, Qiaobo Hao, Huang Liu, Zhiguo Yang, Ji Xie, Ninglun Gu, Jin Yang, Kailai Zhang, Yelun Bao, Jun Wang
DIAG is a Diagnostic Iterative Alignment and Generation framework designed to improve data efficiency in aligning large language models for mathematical reasoning. It adaptively reshapes the practice distribution by first diagnosing valid preference-pair yield to calibrate exploration and exploitation, then generating targeted practice from the model’s failure traces. The approach is theoretically framed as a teacher‑mediated approximation to KL‑regularized reweighting, and experiments show that DIAG increases preference-pair yield and reasoning performance under the same training budget.
By Guhan Chen, Songtao Tian, Bohan Li, Hejin Wang, YeXin Xie, Zixiong Yu
Large language models (LLMs) increasingly improve their reasoning at test time via additional computation, yet most existing works treat each problem in isolation. When problems arrive sequentially, accumulating reusable experience across them can further improve performance.
arXiv:2605. 19723v2 Announce Type: replace-cross Abstract: Mathematical reasoning is essential for problem-solving in education, science, and industry, serving as a crucial benchmark for evaluating artificial intelligence systems.
By Husnain Amjad, Raja Khurram Shahzad, Aamir Shahzad, Mehwish Fatima
YFPO (Yoked Feature Preference Optimization) is a neuron‑guided preference optimization framework that augments standard preference learning with internal neuron‑level rewards. It uses AttnLRP to identify math‑associated internal features and derives an auxiliary reward from the activation margin between preferred and dispreferred responses. Experiments on GSM8K with a compact language model show that these neuron‑guided rewards influence optimization dynamics and yield measurable improvements, indicating that internal representations can serve as lightweight, interpretable signals for reasoning‑oriented post‑training.
By Yifan Le
arXiv:2607. 06974v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly improve their reasoning at test time via additional computation, yet most existing works treat each problem in isolation.
By Ruilin Tong, Dong Gong