arXiv Machine Learning

Optimization as a Dynamical System: Generative Schedules from Latent ODEs

arXiv:2509. 23052v2 Announce Type: replace Abstract: We present a new meta-learning method to determine the optimal learning rate schedule for gradient descent.

arXiv Machine Learning
Sep 11

ExpTest: Loss-Curve Hypothesis Testing for Autonomous Learning-Rate Selection in Deep Neural Networks

ExpTest is an autonomous learning‑rate controller that uses the training loss curve as an online signal to perform sequential statistical tests on theoretically motivated windows, detecting convergent behavior and triggering learning‑rate reductions. It combines a covariance‑based initial learning‑rate estimate, curvature‑motivated window sizing, and a two‑phase test‑driven decay, relying on the approximately exponential decay predicted under linearized network dynamics. Experiments on regression, classification, forecasting, and natural‑language tasks across various architectures show that ExpTest achieves competitive performance compared to hand‑tuned SGD baselines and recent learning‑rate‑free methods, without requiring manual initial learning‑rate selection or predefined scheduling.

By Zan Chaudhry, Naoko Mizuno
arXiv Machine Learning
5d ago

Aurora-X: Built for Extreme Time Series Forecasting

Aurora‑X is a billion‑parameter time‑series foundation model designed for extreme forecasting tasks. It employs a progressive curriculum that starts with channel‑independent pretraining, then adds cross‑variable dependencies, variable context and horizon lengths, and optional future covariates during mid‑training. A variable‑resolution post‑training stage allows adjustable temporal spans per token at inference, while a pattern‑guided mixture‑of‑experts expands capacity through sparse activation and expert specialization. An implicit quantile network head predicts arbitrary quantiles, enhancing probabilistic forecasting flexibility. Experiments on GIFT‑Eval, TIME, FEV‑Bench, TFB, and DAG‑Bench show state‑of‑the‑art performance against both pretrained TSFMs and task‑specific supervised models.

By Xingjian Wu, Chenjuan Guo, Xiangfei Qiu, Zhigang Hu, Hanyin Cheng, Peng Chen, Yang Shu, Jilin Hu, Bin Yang
arXiv Computer Vision
Sep 17

FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation

FlashAR is a lightweight post‑training adaptation framework that converts a pre‑trained raster‑scan autoregressive image model into a highly parallel generator using two‑way next‑token prediction. It preserves the original training objective by keeping the horizontal head for row‑wise prediction and adding a lightweight vertical head for column‑wise prediction, with a learnable fusion gate to combine the two predictions. A two‑stage adaptation pipeline—first initializing the vertical head from the pre‑trained model and then jointly fine‑tuning—yields up to a 22.9× speedup for 512×512 image generation while using only 0.05% of the original training data.

By Junkang Zhou, Yefei He, Feng Chen, Weijie Wang, Bohan Zhuang
arXiv Machine Learning
Jul 8

Drift Happens: An Empirical Study of Neural Architecture Robustness to Temporal Distribution Shift

arXiv:2607. 05908v1 Announce Type: new Abstract: Real-world data distributions evolve over time, inducing temporal distribution shift that can substantially degrade the reliability of deployed machine learning systems.

By Robin Holzinger (Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, USA), Riccardo Colletti (Department of Electrical Engineering and Computer Sciences, University of California, Berkeley, USA)
arXiv AI
Jul 29

CIFNet: An Analytic Neural Learning Framework for Efficient and Calibrated Class-Incremental Learning

arXiv:2509. 11285v2 Announce Type: replace-cross Abstract: Class-Incremental Learning (CIL) in deep neural networks is conventionally framed as an iterative gradient-based optimization problem, incurring high computational cost, hyperparameter sensitivity, and risk of catastrophic forgetting.

By Alejandro Dopico-Castro, Oscar Fontenla-Romero, Bertha Guijarro-Berdi\~nas, Amparo Alonso-Betanzos
arXiv Computer Vision
Aug 27

Efficient Training with Foresight: Multi-Token Auxiliary Supervision for Autoregressive Image Generation

The paper introduces MTAR, a training framework for autoregressive image generation that enhances performance through multi-token prediction, token-level contrastive regularization, and semantic dropping. These components address sparse supervision, improve representation discriminability, and accelerate training without affecting inference. On ImageNet, MTAR outperforms LlamaGen with lower FID and faster training, achieving comparable results in only a third of the iterations.

By Guo Niu, Xiongfei Yao, Teng Wang, Nannan Zhu