arXiv:2506. 14202v4 Announce Type: replace-cross Abstract: End-to-end backpropagation requires storing activations throughout all layers, creating memory bottlenecks that limit model scalability.
By Makoto Shing, Masanori Koyama, Takuya Akiba
arXiv:2606. 10678v1 Announce Type: new Abstract: Transformer-based models have emerged as leading paradigms in time-series forecasting in recent years, employing self-attention mechanisms to capture long-range dependencies.
By Amrijit Biswas, Mustafa Kamal, Robin Krambroeckers, M. M. Lutfe Elahi, Sifat Momen, Nabeel Mohammed, Shafin Rahman
arXiv:2511. 09789v2 Announce Type: replace Abstract: Recent advances in deep forecasting models have achieved remarkable performance, yet most approaches still struggle to provide both accurate predictions and interpretable insights into temporal dynamics.
By Fulong Yao, Wanqing Zhao, Chao Zheng, Xiaofei Han
arXiv:2509. 23052v2 Announce Type: replace Abstract: We present a new meta-learning method to determine the optimal learning rate schedule for gradient descent.
By Matt L. Sampson, Peter Melchior
arXiv:2606. 09658v1 Announce Type: cross Abstract: Muon has recently emerged as a state-of-the-art optimizer for pretraining Large Language Models (LLMs) and vision classifiers.
By Tianyu Ruan, Fengzhuo Zhang, Shuche Wang, Shihua Zhang
arXiv:2605. 10436v2 Announce Type: replace-cross Abstract: Domain Generation Algorithms (DGAs) evolve continuously to evade botnet detection, posing a persistent challenge for dependable network defense.
By Chaeyoung Lee, Chaeri Jung, Seonghoon Jeong