arXiv Machine Learning By Lawrence Feng, Gaurav R. Ghosal, Jacob Mitchell Springer, Ziqian Zhong, Aditi Raghunathan

Early Data Exposure Improves Robustness to Subsequent Fine-Tuning

Read the original on arXiv Machine Learning →

arXiv:2605. 12705v2 Announce Type: replace Abstract: How can we train models whose post-trained capabilities survive subsequent fine-tuning?

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 20

Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining

The paper investigates training-time data augmentation as a regularizer for autoregressive language model pretraining in data‑constrained, compute‑abundant settings. It introduces three orthogonal augmentation categories—token‑level noise, sequence permutations, and target offset prediction—and shows through systematic ablations that each category delays overfitting and reduces validation loss, with random token replacement performing best individually. Combining augmentation categories further lowers the minimum validation loss, demonstrating that such augmentations mitigate data inefficiency in autoregressive pretraining.

By Michael K. Chen, Xikun Zhang, Fan Bai, Zhengding Hu, Zhen Wang