arXiv AI By Shi Feng, Hanlin Zhang, Fan Nie, Sham Kakade, Yiling Chen

Peer-Predictive Self-Training for Language Model Reasoning

Read the original on arXiv AI →

arXiv:2604. 13356v3 Announce Type: replace-cross Abstract: Mechanisms for continued self-improvement of language models without external supervision remain an open challenge.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 12

Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models

The paper introduces SOLID, a framework that enables operations research language models to self-improve without relying on verified answers or external evaluators. SOLID uses solver-generated artifacts from the model’s own rollouts to create pseudo-references, clustering objectives and applying group-relative advantages for dense self-supervision. Experiments on multiple OR benchmarks show that SOLID enhances solution accuracy for both general-purpose and OR-tuned models compared to outcome-only training.

By Rui Zhu, Minglong Cao, Chenyu Zhou, Jianghao Lin, Dongdong Ge