q0: Primitives for Hyper-Epoch Pretraining
arXiv:2606. 03938v1 Announce Type: cross Abstract: Multi-epoch training is becoming the standard now that compute is growing faster than the supply of high-quality text.
arXiv:2606. 03938v1 Announce Type: cross Abstract: Multi-epoch training is becoming the standard now that compute is growing faster than the supply of high-quality text.
arXiv:2609.37076v1 Announce Type: new Abstract: Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this is...
arXiv:2608. 00029v1 Announce Type: cross Abstract: The performance of deep learning models at scale relies heavily on how effectively high-level mathematical operations are mapped to underlying physical hardware.
arXiv:2607. 27692v1 Announce Type: cross Abstract: Top-$K$ sparse attention reduces the cost of Softmax and value aggregation by attending to only a small subset of key--value (KV) entries.
arXiv:2608. 12573v1 Announce Type: new Abstract: Top-k selection is a fundamental computational primitive with applications spanning databases, information retrieval, signal processing, and modern machine learning workloads, including sparse activations and attention pruning.
The paper investigates how attention dynamics evolve across recurrent depth in language models, finding that attention support stabilizes early while hidden states and outputs take longer. It proposes WISE, a training‑free method that uses full attention in early steps and then reuses the discovered sparse working set for later steps, preserving performance on multi‑hop QA tasks. Experiments show that WISE maintains quality up to 2K context, offers measurable speedups, and highlights the importance of recurrent discovery of attention support.
arXiv:2609.37988v1 Announce Type: new Abstract: As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This i...
arXiv:2606. 08854v1 Announce Type: cross Abstract: Standard Reinforcement Learning with Verifiable Rewards (RLVR) training allocates a fixed rollout budget to every query, without regard for what each query's difficulty means for the current policy.
arXiv:2502. 16886v4 Announce Type: replace-cross Abstract: To reduce memory consumption during LLM inference, a handful of methods have been proposed for KV cache pruning.
arXiv:2608. 00916v1 Announce Type: cross Abstract: Modern reranking recipes---billion-scale cross-encoders, mixture-of-experts (MoE) backbones, and distillation against strong teachers---have outpaced the training infrastructure available to most academic groups.
arXiv:2607. 27600v1 Announce Type: new Abstract: Key-value (KV) cache management through compression and eviction strategies has emerged as an important research direction in recent years.
arXiv:2607. 12204v1 Announce Type: new Abstract: Attention can be viewed as an online learner over context, yet existing test-time memories cannot certify that dropping a token leaves outputs unchanged or delete its influence outright.