Fit More and Train Faster With ZeRO via DeepSpeed and FairScale
Related stories
Critical Damping as a Momentum Schedule: Multi-Seed Validation, a Hybrid Recipe, and an Exhaustive Negative Result on Surgical Layer Selection
arXiv:2603. 28921v3 Announce Type: replace-cross Abstract: The critical damping condition of the damped harmonic oscillator model of SGD with momentum (Qian, 1999) yields a momentum schedule with no tuned hyperparameters: mu(t) = 1 - 2*sqrt(alpha(t)).
LESS: Lightweight Evolutionary Supernet Search in Minutes
LESS (Lightweight Evolutionary Supernet Search) is a data‑driven NAS method that uses a brief hard‑path warm‑up and CMA‑ES to evaluate candidate architectures as decoded hard genotypes after six supernet updates. On NAS‑Bench‑201, LESS attains 93.189 % CIFAR‑10 accuracy in just 409.1 seconds, nearly matching FairNAS while using only about 1/24 of its search time. The approach also transfers well to CIFAR‑100, ImageNet16‑120, and the larger DARTS space, achieving high accuracies with searches completed in roughly 43.5 minutes on a single GPU.
Depth and Scale in the Sub-150M Regime: JugnuLM-53M vs JugnuLM-110M
arXiv:2609.14715v1 Announce Type: new Abstract: We scale our conventional sub-150M pretraining recipe from 53.5M to 109.7M parameters, holding the method fixed (Qwen3-style decoder with grouped-query...
Miles v0.1: Production-Level Post-Training
arXiv:2609.08368v1 Announce Type: new Abstract: We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each st...
Many Optimizers But Only One Training Path: Repeated Resampling for Adaptive Optimizer Selection
The paper introduces Repeated Optimizer Resampling (ROR), a method that treats optimizer choice as a hyperparameter and searches for the best optimizer during a single training run. ROR periodically scouts each candidate optimizer for a short number of epochs, then continues training with the best scout, allowing the optimizer to change over time. Experiments on MNIST, Fashion‑MNIST, and motor insurance claim‑count models show that one‑epoch ROR uses only 24–35% of the training effort required to exhaustively evaluate all optimizers while achieving comparable performance.
M+Adam: Low-Precision Training via Additive-Multiplicative Optimization
arXiv:2607. 10611v1 Announce Type: new Abstract: Training with quantized weights can reduce costs but often results in degraded accuracy, especially when optimization is carried out in low precision, without storing high-precision copies.
Adaptive Runge-Kutta Step Control Buys Training Loss, Not Generalization: An Honest Compute-Matched Study of RK-Adam Optimizers
arXiv:2607. 14516v1 Announce Type: new Abstract: Interpreting optimizers as gradient-flow discretizations has motivated applying higher-order Runge-Kutta (RK) integrators to neural networks.
Wiring Beats Blending: Structure-Aware Compensation for Transformer Downscaling
The paper investigates converting a large pretrained transformer (1.4 B parameters) into a smaller sibling (410 M) by studying representation alignment and parameter projection. It finds that dense weight projection destroys structure, and that a low‑budget, structure‑aware compensation—separating least‑squares function alignment from variance‑preserving rescaling—yields significant gains on token‑efficient training, outperforming subcloning and standard distillation pipelines at matched budgets.
Velocity Scheduled Flow Matching
Flow matching trains a neural network to regress the conditional velocity along a linear interpolant between noise and data, and the number of network evaluations~(NFE) sets the cost of sampling. The straight-line interpolant carries an implicit choice: the sample moves at constant speed throughout the trajectory.
Difficulty-Calibrated Interpolation Paths for Conditional Flow Matching
The paper introduces Difficulty-Calibrated Flow Matching, a method that adapts the noise-to-data interpolation schedule in Conditional Flow Matching based on a pilot run’s loss profile. By setting the schedule to the quantile function of this difficulty profile, the training trajectory spends more time where the velocity is hardest to learn. Experiments on CIFAR-10, MNIST, and Fashion‑MNIST show that this calibrated path achieves the best FID on CIFAR‑10 and outperforms all fixed schedules in large‑batch, few‑update settings, where compute is most limited.
AQLoRA: A Zero-Search Recipe for Fast Quantized LoRA Fine-Tuning
arXiv:2608.23816v1 Announce Type: new Abstract: Quantized fine-tuning (QLoRA) saves memory but not time. It dequantizes every 4-bit weight on the fly, so it trains more slowly than fp16 LoRA. We pres...