q0: Primitives for Hyper-Epoch Pretraining
arXiv:2606. 03938v1 Announce Type: cross Abstract: Multi-epoch training is becoming the standard now that compute is growing faster than the supply of high-quality text.
arXiv:2606. 08578v1 Announce Type: new Abstract: Recently, large time series models (LTSMs) have gained increasing attention due to their similarities to large language models, including flexible context length, scalability, and task generality, outperforming advanced task-specific models.
arXiv:2606. 03938v1 Announce Type: cross Abstract: Multi-epoch training is becoming the standard now that compute is growing faster than the supply of high-quality text.
arXiv:2509. 22020v2 Announce Type: replace Abstract: While recent advances in machine learning have equipped Weather Foundation Models (WFMs) with substantial generalization capabilities across diverse downstream tasks, the escalating computational requirements associated with their expanding scale increasingly hinder practical deployment.
The paper investigates how the local landscape geometry of language model pre‑training evolves, identifying two distinct phases. In Phase I, the landscape starts sharp, causing instability and loss plateaus at high learning rates, which explains the need for learning‑rate warmup and suggests longer warmups for larger peak rates. In Phase II, the geometry is governed by gradient noise scale, revealing a depth‑flatness trade‑off that motivates a dynamic batch‑size scheduler that starts small and grows later in training.
arXiv:2606. 10466v1 Announce Type: cross Abstract: In time-series generation, existing approaches typically handcraft ortrain a separate model for each dataset, which hinders their scalability and fails to leverage shared temporal structures across domains.
arXiv:2605.07111v3 Announce Type: replace-cross Abstract: Recent literature on fine-tuning Large Language Models highlights a fundamental debate. While Full Fine-Tuning (FFT) provides greater represe...
The paper investigates training-time data augmentation as a regularizer for autoregressive language model pretraining in data‑constrained, compute‑abundant settings. It introduces three orthogonal augmentation categories—token‑level noise, sequence permutations, and target offset prediction—and shows through systematic ablations that each category delays overfitting and reduces validation loss, with random token replacement performing best individually. Combining augmentation categories further lowers the minimum validation loss, demonstrating that such augmentations mitigate data inefficiency in autoregressive pretraining.
arXiv:2609.37169v1 Announce Type: cross Abstract: Mid-training equips pretrained large language models with specialized and reasoning capabilities, but the returns of this stage are bounded since add...
arXiv:2502. 11034v3 Announce Type: replace Abstract: Loss spikes remain a persistent obstacle in large-scale language model pretraining.
arXiv:2609.37076v1 Announce Type: new Abstract: Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this is...
arXiv:2608. 13262v1 Announce Type: cross Abstract: Time series foundation models (TSFMs) have advanced primarily through architectural innovation, while training regimes for large-scale heterogeneous corpora remain under-explored.
arXiv:2606. 16246v1 Announce Type: cross Abstract: As AI labs approach a data ceiling where compute capacity outpaces the rate of new high-quality text generation, language model pretraining is shifting toward a data-constrained, compute-abundant regime that demands productive multi-epoch training on fixed corpora.
arXiv:2607. 22769v1 Announce Type: cross Abstract: The training efficacy of large language models (LLMs) is fundamentally constrained by the quality and composition of training data.