arXiv Computation and Language By Shuai Wang, Jiayi Kuang, Yinghui Li, Haojing Huang, Xinnian Liang, Ying Shen, Liang Lin

INSPIRE: An Internalize-Then-Improve Approach for Example-Driven Mathematical Reasoning

Read the original on arXiv Computation and Language →

The paper introduces INSPIRE, an Internalize-Then-Improve framework designed to enhance example-driven mathematical reasoning in large language models. It combines Reference-Guided Student Internalization (RGSI) to generate high-quality preference pairs with a staged rubric preference training that separates method learning from correctness. Experiments across various model sizes show consistent gains, even outperforming larger open-source models, and maintain performance on out-of-distribution mathematical tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jun 30

Beyond Scaling Law: A Data-Efficient Distillation Framework for Reasoning

arXiv:2508. 09883v2 Announce Type: replace-cross Abstract: Large language models (LLMs) demonstrate remarkable reasoning capabilities in tasks such as algorithmic coding and mathematical problem-solving.

By Xiaojun Wu, Xiaoguang Jiang, Huiyang Li, Jucai Zhai, Dengfeng Liu, Qiaobo Hao, Huang Liu, Zhiguo Yang, Ji Xie, Ninglun Gu, Jin Yang, Kailai Zhang, Yelun Bao, Jun Wang
arXiv Computation and Language
Aug 25

DIAG: Diagnostic Iterative Alignment and Generation for Data-Efficient Mathematical Preference Distillation

DIAG is a Diagnostic Iterative Alignment and Generation framework designed to improve data efficiency in aligning large language models for mathematical reasoning. It adaptively reshapes the practice distribution by first diagnosing valid preference-pair yield to calibrate exploration and exploitation, then generating targeted practice from the model’s failure traces. The approach is theoretically framed as a teacher‑mediated approximation to KL‑regularized reweighting, and experiments show that DIAG increases preference-pair yield and reasoning performance under the same training budget.

By Guhan Chen, Songtao Tian, Bohan Li, Hejin Wang, YeXin Xie, Zixiong Yu
arXiv Computation and Language
Sep 16

YFPO: Yoked Feature Preference Optimization with Neuron-Guided Rewards

YFPO (Yoked Feature Preference Optimization) is a neuron‑guided preference optimization framework that augments standard preference learning with internal neuron‑level rewards. It uses AttnLRP to identify math‑associated internal features and derives an auxiliary reward from the activation margin between preferred and dispreferred responses. Experiments on GSM8K with a compact language model show that these neuron‑guided rewards influence optimization dynamics and yield measurable improvements, indicating that internal representations can serve as lightweight, interpretable signals for reasoning‑oriented post‑training.

By Yifan Le