arXiv:2602. 09530v2 Announce Type: replace-cross Abstract: We introduce AutoSpec, a neural network framework for discovering iterative spectral algorithms for large-scale numerical linear algebra and numerical optimization.
By Zihang Liu, Oleg Balabanov, Yaoqing Yang, Michael W. Mahoney
arXiv:2607. 02050v1 Announce Type: new Abstract: Motivated by the challenge of stabilizing a general unknown linear dynamical system (LDS) from observations, we study the natural prerequisite of online prediction.
By Yuval Ran-Milo, Angelos Assos, Elad Hazan
arXiv:2604. 25021v2 Announce Type: replace Abstract: We study online regression with the square loss in a reproducing kernel Hilbert space under a dynamic regret criterion.
By Dmitry B. Rokhlin, Georgiy A. Karapetyants
arXiv:2608. 14691v1 Announce Type: new Abstract: Sequence models are conventionally distinguished by their backbone, the mechanism that routes information across positions, such as attention or recurrence.
By Ahmed Nebli, Hadi Saadatdoorabi, Christopher Keibel, Kevin Yam
arXiv:2310. 15976v4 Announce Type: replace Abstract: signSGD is attractive in nonconvex optimization because it communicates sign-valued rather than full-precision gradients.
By Zhen Qin, Zhishuai Liu, Pan Xu
arXiv:2606. 21253v2 Announce Type: replace Abstract: Continual learning that is gradient-free, local, online, and append-only is attractive for edge and streaming deployment, but its value is usually argued informally.
By Jianwei Lou (RailMind Systems, Neuss, Germany)
arXiv:2604. 07328v3 Announce Type: replace Abstract: How does the choice of training data influence an AI model?
By Sam Gunn
arXiv:2608. 11019v1 Announce Type: new Abstract: Modeling spatiotemporal dynamical systems governed by partial differential equations (PDEs) poses two major challenges: it either requires expensive physics-based simulators that entail iterative numerical solving at high computational cost, or it depends on abundant training data, yet purely data-driven models often generalize poorly to downstream dynamic operating conditions.
By Hengbo Xiao, Jiale Liu, Jiahao Song, Guannan He
arXiv:2606. 04031v1 Announce Type: new Abstract: Coupled gradient descent--where the update of one parameter block depends on another--underlies bilevel optimization, two-time-scale stochastic approximation, and adversarial training.
By Ahanaf Hasan Ariq
arXiv:2606. 04834v1 Announce Type: new Abstract: Minimum Description Length (MDL) formalizes the principle of Occam's razor by optimizing the total description length: $L(\mathrm{model})+L(\mathrm{data} \ | \ \mathrm{model})$.
By Qian Li, Xinyu Mao, Shang-Hua Teng, Guangxu Yang
arXiv:2602. 16864v2 Announce Type: replace-cross Abstract: Time series (TS) modeling has come a long way from early statistical, mainly linear, approaches to the current trend in TS foundation models.
By Daniel Durstewitz, Christoph J\"urgen Hemmer, Florian Hess, Charlotte Ricarda Doll, Lukas Eisenmann
arXiv:2607. 20594v1 Announce Type: cross Abstract: When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm?
By Tong Zhang, Junhao Hu, Yun Peng, Tao Xie