arXiv Computation and Language

YFPO: Yoked Feature Preference Optimization with Neuron-Guided Rewards

YFPO (Yoked Feature Preference Optimization) is a neuron‑guided preference optimization framework that augments standard preference learning with internal neuron‑level rewards. It uses AttnLRP to identify math‑associated internal features and derives an auxiliary reward from the activation margin between preferred and dispreferred responses. Experiments on GSM8K with a compact language model show that these neuron‑guided rewards influence optimization dynamics and yield measurable improvements, indicating that internal representations can serve as lightweight, interpretable signals for reasoning‑oriented post‑training.

Hugging Face Trending Papers
Jun 24

MiniOpt: Reasoning to Model and Solve General Optimization Problems with Limited Resources

Achieving strong optimization generalization across diverse optimization problems while requiring limited training resources remains a challenging problem for optimization-oriented large language models (LLMs). Existing approaches typically rely on large-scale supervised datasets, costly reasoning annotations, and expensive intermediate step verification, resulting in substantial training overhead.

arXiv Computation and Language
Aug 31

INSPIRE: An Internalize-Then-Improve Approach for Example-Driven Mathematical Reasoning

The paper introduces INSPIRE, an Internalize-Then-Improve framework designed to enhance example-driven mathematical reasoning in large language models. It combines Reference-Guided Student Internalization (RGSI) to generate high-quality preference pairs with a staged rubric preference training that separates method learning from correctness. Experiments across various model sizes show consistent gains, even outperforming larger open-source models, and maintain performance on out-of-distribution mathematical tasks.

By Shuai Wang, Jiayi Kuang, Yinghui Li, Haojing Huang, Xinnian Liang, Ying Shen, Liang Lin
Hugging Face Trending Papers
Jul 2

Neuron-Aware Data Selection for Annotation-Free LLM Self-Distillation

Post-training large language models (LLMs) without real-world interaction feedback or human-labeled supervision remains challenging, particularly in specialized domains where expert annotations are costly to obtain. Recent annotation-free self-evolution methods address this by using the model's own outputs as supervision signals, constructing a teacher via additional context and aggregating predictions across multiple rollouts through majority voting to produce pseudo-labels.

arXiv Machine Learning
Jun 11

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal

arXiv:2606. 12360v1 Announce Type: new Abstract: Language-model post-training is the main stage at which model behavior is shaped, yet it still largely involves optimization of scalar rewards that summarize diverse desiderata.

By Leon Bergen, Usha Bhalla, Sidharth Baskaran, Max Loeffler, Raphael Sarfati, Dhruvil Gala, Ryan Panwar, Santiago Aranguri, Thomas Fel, Atticus Geiger, Matthew Kowal, Siddharth Boppana, Daniel Balsam, Owen Lewis, Jack Merullo, Thomas McGrath, Ekdeep Singh Lubana