arXiv Computation and Language

Strategically Diverse Sampling for Self-Training

The paper introduces two new sampling techniques—GROOT and Verbalized Sampling—to create strategically diverse training data for self-training large language models. By focusing on substantive variation in problem-solving approaches rather than just correctness, the authors demonstrate that models trained on this data outperform those trained on IID samples across competitive programming and Next‑Chapter Prediction tasks. Notably, self‑training with strategically diverse but incorrect traces from Qwen3‑4B surpasses IID distillation from a 235B teacher, challenging assumptions about the importance of correctness and teacher scale.

arXiv Machine Learning
Aug 11

On the Effect of Sampling Diversity in Scaling LLM Inference

arXiv:2502. 11027v5 Announce Type: replace Abstract: Large language model (LLM) scaling inference is key to unlocking greater performance, and leveraging diversity has proven an effective way to enhance it.

By Tianchun Wang, Zichuan Liu, Yuanzhou Chen, Jonathan Light, Weiyang Liu, Haifeng Chen, Xiang Zhang, Wei Cheng
arXiv AI
Jul 16

Representation-Based Exploration for Language Models: From Test-Time to Post-Training

arXiv:2510. 11686v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) promises to expand the capabilities of language models, but it is unclear if current RL techniques promote the discovery of novel behaviors, or simply sharpen those already present in the base model.

By Jens Tuyls, Dylan J. Foster, Akshay Krishnamurthy, Jordan T. Ash
arXiv Computation and Language
Sep 23

Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning

The paper proposes a new approach to large language model (LLM) reasoning that moves beyond naive repeated sampling. Instead of generating many independent solutions, it first samples problem‑specific concepts, hints, or strategies and conditions answer generation on them, producing a single trajectory of diverse concepts. A small concept generator is then trained via reinforcement learning to maximize downstream success, leading to significant improvements in pass@k on hard mathematical reasoning tasks compared to both naive sampling and concepts from larger untuned models, and the trained generator transfers to unseen answer generators, including those from different model families.

By Ismail Labiad, Matthieu Kowalski, Marc Schoenauer, R\'emi Munos, Julia Kempe
arXiv Computation and Language
Sep 2

From Rollouts to Recipes: Self-Contained Post-Training for LLMs

The paper introduces Self‑Routing, a post‑training framework that tailors optimization for each sample based on its rollout correctness and confidence. Instead of applying a single recipe to all data, samples are routed to different strategies—GRPO, on‑policy self‑distillation, regularization, or skipped—allowing training to adapt without external teachers or extra annotations. Experiments on Qwen3 and Qwen3.5 show consistent improvements over uniform methods and reveal that the routing distribution evolves during training, reducing unnecessary updates on low‑signal or already stable samples.

By Yifei Li, Lingling Zhang, Muye Huang, Zihan Ma, Jiashuai Liu, Jun Liu