arXiv Computation and Language

INSPIRE: An Internalize-Then-Improve Approach for Example-Driven Mathematical Reasoning

The paper introduces INSPIRE, an Internalize-Then-Improve framework designed to enhance example-driven mathematical reasoning in large language models. It combines Reference-Guided Student Internalization (RGSI) to generate high-quality preference pairs with a staged rubric preference training that separates method learning from correctness. Experiments across various model sizes show consistent gains, even outperforming larger open-source models, and maintain performance on out-of-distribution mathematical tasks.

arXiv AI
Jun 30

Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning

arXiv:2508. 09883v2 Announce Type: replace-cross Abstract: Large language models (LLMs) demonstrate remarkable reasoning capabilities in tasks such as algorithmic coding and mathematical problem-solving.

By Xiaojun Wu, Xiaoguang Jiang, Huiyang Li, Jucai Zhai, Dengfeng Liu, Qiaobo Hao, Huang Liu, Zhiguo Yang, Ji Xie, Ninglun Gu, Jin Yang, Kailai Zhang, Yelun Bao, Jun Wang
arXiv Computation and Language
Aug 25

DIAG: Diagnostic Iterative Alignment and Generation for Data-Efficient Mathematical Preference Distillation

DIAG is a Diagnostic Iterative Alignment and Generation framework designed to improve data efficiency in aligning large language models for mathematical reasoning. It adaptively reshapes the practice distribution by first diagnosing valid preference-pair yield to calibrate exploration and exploitation, then generating targeted practice from the model’s failure traces. The approach is theoretically framed as a teacher‑mediated approximation to KL‑regularized reweighting, and experiments show that DIAG increases preference-pair yield and reasoning performance under the same training budget.

By Guhan Chen, Songtao Tian, Bohan Li, Hejin Wang, YeXin Xie, Zixiong Yu
arXiv Computation and Language
Sep 16

YFPO: Yoked Feature Preference Optimization with Neuron-Guided Rewards

YFPO (Yoked Feature Preference Optimization) is a neuron‑guided preference optimization framework that augments standard preference learning with internal neuron‑level rewards. It uses AttnLRP to identify math‑associated internal features and derives an auxiliary reward from the activation margin between preferred and dispreferred responses. Experiments on GSM8K with a compact language model show that these neuron‑guided rewards influence optimization dynamics and yield measurable improvements, indicating that internal representations can serve as lightweight, interpretable signals for reasoning‑oriented post‑training.

By Yifan Le
arXiv Machine Learning
Aug 28

Learning to Reason with Curriculum I: Provable Benefits of Autocurriculum

The paper investigates whether the high costs of training chain-of-thought reasoning models can be reduced through algorithmic design. It introduces an autocurriculum approach that lets the model select which problems to focus on during training, showing that this method provably improves both supervised fine‑tuning and reinforcement learning. For supervised fine‑tuning, autocurriculum requires exponentially fewer reasoning demonstrations by targeting prompts where the model struggles, while for reinforcement learning it decouples computational cost from the quality of the reference model, making the burn‑in cost nearly independent of target accuracy.

By Nived Rajaraman, Audrey Huang, Miro Dudik, Robert Schapire, Dylan J. Foster, Akshay Krishnamurthy
arXiv Machine Learning
1d ago

The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models

arXiv:2610. 02191v1 Announce Type: new Abstract: While Large Language Models (LLMs) have demonstrated striking capabilities on frontier mathematical problems, it remains unclear whether they possess the structural mathematical understanding underlying their solutions.

By Shuo Xing, Zilin Dai, Chengyuan Qian, Fangzhou Lin, Wenjing Chen, Ping He, Pan Lu, Alvaro Velasquez, Mohit Bansal, Zhengzhong Tu