arXiv Machine Learning By Matt L. Sampson, Peter Melchior

Optimization as a Dynamical System: Generative Schedules from Latent ODEs

Read the original on arXiv Machine Learning →

arXiv:2509. 23052v2 Announce Type: replace Abstract: We present a new meta-learning method to determine the optimal learning rate schedule for gradient descent.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 11

ExpTest: Loss-Curve Hypothesis Testing for Autonomous Learning-Rate Selection in Deep Neural Networks

ExpTest is an autonomous learning‑rate controller that uses the training loss curve as an online signal to perform sequential statistical tests on theoretically motivated windows, detecting convergent behavior and triggering learning‑rate reductions. It combines a covariance‑based initial learning‑rate estimate, curvature‑motivated window sizing, and a two‑phase test‑driven decay, relying on the approximately exponential decay predicted under linearized network dynamics. Experiments on regression, classification, forecasting, and natural‑language tasks across various architectures show that ExpTest achieves competitive performance compared to hand‑tuned SGD baselines and recent learning‑rate‑free methods, without requiring manual initial learning‑rate selection or predefined scheduling.

By Zan Chaudhry, Naoko Mizuno
arXiv Machine Learning
5d ago

Aurora-X: Built for Extreme Time Series Forecasting

Aurora‑X is a billion‑parameter time‑series foundation model designed for extreme forecasting tasks. It employs a progressive curriculum that starts with channel‑independent pretraining, then adds cross‑variable dependencies, variable context and horizon lengths, and optional future covariates during mid‑training. A variable‑resolution post‑training stage allows adjustable temporal spans per token at inference, while a pattern‑guided mixture‑of‑experts expands capacity through sparse activation and expert specialization. An implicit quantile network head predicts arbitrary quantiles, enhancing probabilistic forecasting flexibility. Experiments on GIFT‑Eval, TIME, FEV‑Bench, TFB, and DAG‑Bench show state‑of‑the‑art performance against both pretrained TSFMs and task‑specific supervised models.

By Xingjian Wu, Chenjuan Guo, Xiangfei Qiu, Zhigang Hu, Hanyin Cheng, Peng Chen, Yang Shu, Jilin Hu, Bin Yang
arXiv Computer Vision
Sep 17

FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation

FlashAR is a lightweight post‑training adaptation framework that converts a pre‑trained raster‑scan autoregressive image model into a highly parallel generator using two‑way next‑token prediction. It preserves the original training objective by keeping the horizontal head for row‑wise prediction and adding a lightweight vertical head for column‑wise prediction, with a learnable fusion gate to combine the two predictions. A two‑stage adaptation pipeline—first initializing the vertical head from the pre‑trained model and then jointly fine‑tuning—yields up to a 22.9× speedup for 512×512 image generation while using only 0.05% of the original training data.

By Junkang Zhou, Yefei He, Feng Chen, Weijie Wang, Bohan Zhuang