arXiv Machine Learning By Jiaming Li, Haoran Ye, Yukun Chen, Xinyue Li, Lei Zhang, Hamid Alinejad-Rokny, Jimmy Chih-Hsien Peng, Min Yang

Breaking the Block: Preserving Data Continuity to Train Superior SAEs for Instruct Models

Read the original on arXiv Machine Learning →

arXiv:2506. 07691v2 Announce Type: replace-cross Abstract: Sparse Autoencoders (SAEs) are a cornerstone of mechanistic interpretability.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 16

Data Augmentations for Data-Constrained Language Model Pretraining

arXiv:2606. 16246v1 Announce Type: cross Abstract: As AI labs approach a data ceiling where compute capacity outpaces the rate of new high-quality text generation, language model pretraining is shifting toward a data-constrained, compute-abundant regime that demands productive multi-epoch training on fixed corpora.

By Michael K. Chen, Xikun Zhang, Zhen Wang