arXiv Machine Learning

Early Data Exposure Improves Robustness to Subsequent Fine-Tuning

arXiv:2605. 12705v2 Announce Type: replace Abstract: How can we train models whose post-trained capabilities survive subsequent fine-tuning?

arXiv AI
Aug 20

Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining

The paper investigates training-time data augmentation as a regularizer for autoregressive language model pretraining in data‑constrained, compute‑abundant settings. It introduces three orthogonal augmentation categories—token‑level noise, sequence permutations, and target offset prediction—and shows through systematic ablations that each category delays overfitting and reduces validation loss, with random token replacement performing best individually. Combining augmentation categories further lowers the minimum validation loss, demonstrating that such augmentations mitigate data inefficiency in autoregressive pretraining.

By Michael K. Chen, Xikun Zhang, Fan Bai, Zhengding Hu, Zhen Wang
arXiv AI
Jun 16

Data Augmentations for Data-Constrained Language Model Pretraining

arXiv:2606. 16246v1 Announce Type: cross Abstract: As AI labs approach a data ceiling where compute capacity outpaces the rate of new high-quality text generation, language model pretraining is shifting toward a data-constrained, compute-abundant regime that demands productive multi-epoch training on fixed corpora.

By Michael K. Chen, Xikun Zhang, Zhen Wang
arXiv Machine Learning
Sep 16

Test-Time Unlearning via Sparse Autoencoder

arXiv:2609.16229v1 Announce Type: new Abstract: Machine unlearning aims to remove specific knowledge from a trained large language model (LLM) without retraining from scratch. Existing methods modify...

By Pingzhi Li, Jinhao Duan, Vaishnav Tadiparthi, Nakul Agarwal, Kwonjoon Lee, Ehsan Moradi Pari, Hossein Nourkhiz Mahjoub, Sijia Liu, Tianlong Chen