arXiv Computation and Language By Alexander Gurung, Esmeralda S. Whitammer, Mirella Lapata

Strategically Diverse Sampling for Self-Training

Read the original on arXiv Computation and Language →

The paper introduces two new sampling techniques—GROOT and Verbalized Sampling—to create strategically diverse training data for self-training large language models. By focusing on substantive variation in problem-solving approaches rather than just correctness, the authors demonstrate that models trained on this data outperform those trained on IID samples across competitive programming and Next‑Chapter Prediction tasks. Notably, self‑training with strategically diverse but incorrect traces from Qwen3‑4B surpasses IID distillation from a 235B teacher, challenging assumptions about the importance of correctness and teacher scale.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Aug 11

On the Effect of Sampling Diversity in Scaling LLM Inference

arXiv:2502. 11027v5 Announce Type: replace Abstract: Large language model (LLM) scaling inference is key to unlocking greater performance, and leveraging diversity has proven an effective way to enhance it.

By Tianchun Wang, Zichuan Liu, Yuanzhou Chen, Jonathan Light, Weiyang Liu, Haifeng Chen, Xiang Zhang, Wei Cheng
arXiv AI
Jul 16

Representation-Based Exploration for Language Models: From Test-Time to Post-Training

arXiv:2510. 11686v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) promises to expand the capabilities of language models, but it is unclear if current RL techniques promote the discovery of novel behaviors, or simply sharpen those already present in the base model.

By Jens Tuyls, Dylan J. Foster, Akshay Krishnamurthy, Jordan T. Ash
arXiv Computation and Language
Sep 23

Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning

The paper proposes a new approach to large language model (LLM) reasoning that moves beyond naive repeated sampling. Instead of generating many independent solutions, it first samples problem‑specific concepts, hints, or strategies and conditions answer generation on them, producing a single trajectory of diverse concepts. A small concept generator is then trained via reinforcement learning to maximize downstream success, leading to significant improvements in pass@k on hard mathematical reasoning tasks compared to both naive sampling and concepts from larger untuned models, and the trained generator transfers to unseen answer generators, including those from different model families.

By Ismail Labiad, Matthieu Kowalski, Marc Schoenauer, R\'emi Munos, Julia Kempe