Hugging Face Blog

Fit More and Train Faster With ZeRO via DeepSpeed and FairScale

arXiv Machine Learning
3d ago

LESS: Lightweight Evolutionary Supernet Search in Minutes

LESS (Lightweight Evolutionary Supernet Search) is a data‑driven NAS method that uses a brief hard‑path warm‑up and CMA‑ES to evaluate candidate architectures as decoded hard genotypes after six supernet updates. On NAS‑Bench‑201, LESS attains 93.189 % CIFAR‑10 accuracy in just 409.1 seconds, nearly matching FairNAS while using only about 1/24 of its search time. The approach also transfers well to CIFAR‑100, ImageNet16‑120, and the larger DARTS space, achieving high accuracies with searches completed in roughly 43.5 minutes on a single GPU.

By Aviral Gandhi, Jinglue Xu, Jialong Li, Hitoshi Iba
arXiv Machine Learning
Sep 10

Miles v0.1: Production-Level Post-Training

arXiv:2609.08368v1 Announce Type: new Abstract: We present Miles v0.1, a full-stack, production-ready system for frontier post-training. Building upon the clean design of slime, Miles designs each st...

By RadixArk, :, Tom Chen, Mao Cheng, Shi Dong, Kangrui Du, Yanbin Jiang, Jiajun Li, Yiming Li, Tao Lin, Yusheng Su, Andy Ye, Yueming Yuan, Zhichen Zeng
arXiv Machine Learning
Aug 20

Many Optimizers But Only One Training Path: Repeated Resampling for Adaptive Optimizer Selection

The paper introduces Repeated Optimizer Resampling (ROR), a method that treats optimizer choice as a hyperparameter and searches for the best optimizer during a single training run. ROR periodically scouts each candidate optimizer for a short number of epochs, then continues training with the best scout, allowing the optimizer to change over time. Experiments on MNIST, Fashion‑MNIST, and motor insurance claim‑count models show that one‑epoch ROR uses only 24–35% of the training effort required to exhaustively evaluate all optimizers while achieving comparable performance.

By Ronald Richman, Mario V. W\"uthrich
arXiv AI
Sep 25

Wiring Beats Blending: Structure-Aware Compensation for Transformer Downscaling

The paper investigates converting a large pretrained transformer (1.4 B parameters) into a smaller sibling (410 M) by studying representation alignment and parameter projection. It finds that dense weight projection destroys structure, and that a low‑budget, structure‑aware compensation—separating least‑squares function alignment from variance‑preserving rescaling—yields significant gains on token‑efficient training, outperforming subcloning and standard distillation pipelines at matched budgets.

By Ravi Satya Durga Prasad Yenugula
Hugging Face Trending Papers
Jul 13

Velocity Scheduled Flow Matching

Flow matching trains a neural network to regress the conditional velocity along a linear interpolant between noise and data, and the number of network evaluations~(NFE) sets the cost of sampling. The straight-line interpolant carries an implicit choice: the sample moves at constant speed throughout the trajectory.

arXiv Computer Vision
Aug 24

Difficulty-Calibrated Interpolation Paths for Conditional Flow Matching

The paper introduces Difficulty-Calibrated Flow Matching, a method that adapts the noise-to-data interpolation schedule in Conditional Flow Matching based on a pilot run’s loss profile. By setting the schedule to the quantile function of this difficulty profile, the training trajectory spends more time where the velocity is hardest to learn. Experiments on CIFAR-10, MNIST, and Fashion‑MNIST show that this calibrated path achieves the best FID on CIFAR‑10 and outperforms all fixed schedules in large‑batch, few‑update settings, where compute is most limited.

By Airin Akter Tania, Md Raihan Khan