arXiv AI By Jingwen Liu, Ezra Edelman, Surbhi Goel, Bingbin Liu

Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases

Read the original on arXiv AI →

The paper titled "Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases" examines how training on a smaller dataset with repeated samples can reduce computational cost compared to using a larger dataset. Across various tasks, architectures, and optimizers, this phenomenon cannot be explained by existing theory. The authors attribute the speedup to layer‑wise growth driven by sampling biases, which is more pronounced with smaller datasets, and provide both theoretical analysis and empirical evidence to support this claim.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

What Matters in On-Policy Distillation? A Perspective on Data Efficiency and Data Selection

The paper investigates data efficiency and selection in On‑Policy Distillation (OPD) for large language models. It shows that 1‑shot OPD—training on a single example—consistently improves performance, especially when the example is hard, and that longer chain‑of‑thought (CoT) paths drive the gains rather than token entropy. Based on these findings, the authors propose a simple hard‑example selection strategy that, using only eight carefully chosen hard examples, matches the performance of a 17,000‑example baseline across models from 1.5B to 7B parameters.

By Zhinan Hou, Jiaqi Zhang, Xunliang Cai, Keyou You
arXiv AI
Sep 25

To Think or Not to Think: Allocating Reasoning Where It Helps

The paper introduces CARE, a contrastive accuracy reward estimation method that adaptively adjusts reasoning length for large language models. By comparing beneficial length adjustments from online sampled responses, CARE applies adaptive length rewards within Group Relative Policy Optimization without extra hyperparameters or inference cost. Experiments on multiple reasoning benchmarks show that CARE improves Pass@1 by up to 4% while reducing reasoning length by 37%, achieving higher token efficiency.

By Zhengdong He, Yunfan Zhou, Jianguo Yao, Haibing Guan, Xijun Li
arXiv Machine Learning
Aug 28

Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO

The paper investigates Evolution Strategies (ES) as a memory‑efficient post‑training method for large language model (LLM) reasoning. It demonstrates that ES outperforms Group Relative Policy Optimization (GRPO) by achieving broader reasoning coverage, improving Pass@K metrics, and avoiding entropy collapse. The study also reveals that ES’s performance gains stem from sparse, high‑magnitude parameter updates, do not cause catastrophic forgetting, and can be combined with GRPO in a sequential training strategy.

By Yunpeng Ba, Zhi Zheng, Yue Xie, Jiaqing Li, Xialiang Tong, Tao Zhong, Mingxuan Yuan, Zhichao Lu, Xuyang Wu, Zhenkun Wang
arXiv AI
6d ago

Learning to Stop without Learning to Stop: Self-Supervised Confidence Training Improves Reasoning Efficiency

The paper demonstrates that fine‑tuning reasoning models to predict their own confidence at intermediate steps—using only 600 self‑supervised examples—substantially improves inference efficiency. Without adding any explicit stopping or length penalties, the models generate up to 25 % fewer tokens while maintaining accuracy on mathematical, scientific, and coding benchmarks across several architectures. The study finds that confidence supervision preserves the models’ high‑level reasoning structure rather than merely suppressing specific behaviors.

By Parsa Hosseini, Akasha Tigalappanavara, Sumit Nawathe, Chenrui Fan, Sourya Basu, Genta Indra Winata, Anirban Das, Soheil Feizi, Nima Chitsazan