arXiv:2608. 06597v1 Announce Type: cross Abstract: A scientific theory of deep learning, comprising learning dynamics and statistical properties of learned models, is rapidly gaining attention.
By Bj\"orn Ladewig, Ibrahim Talha Ersoy, Karoline Wiesner
arXiv:2508. 12448v2 Announce Type: replace-cross Abstract: In-context learning (ICL) lets large language models (LLMs) solve new tasks from prompts alone, across an ever-widening range of domains, yet the mechanisms underlying this ability remain poorly understood.
By Yeongwoo Song, Jaeyong Bae, Dong-Kyum Kim, Hawoong Jeong
arXiv:2608. 01833v1 Announce Type: cross Abstract: Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generalization.
By Lai Shun Chan, Xiaotian Zhang, Yue Shang, Ge Zhang, Entao Yang
arXiv:2501. 02436v5 Announce Type: replace Abstract: Advancements in artificial intelligence call for a deeper understanding of the fundamental mechanisms underlying deep learning.
By Yuchen Lin, Yong Zhang, Sihan Feng, Hong Zhao
arXiv:2608. 09396v1 Announce Type: new Abstract: Invariant learning seeks representations that remain predictive across environments, yet the behavior of its objectives along the regularization path is often opaque.
By Pinli Wang, Yue He, Peng Cui
arXiv:2607. 07763v1 Announce Type: new Abstract: World models are typically trained to predict discrete-time physical dynamics with a fixed step size baked into the model weights, preventing prediction at variable temporal resolutions.
By Eli Laird, Corey Clark
arXiv:2606. 31282v1 Announce Type: new Abstract: Modern deep neural networks often contain far more parameters than needed to fit their training data, yet they achieve impressive generalization.
By Ari Pakman, Lior Kreimer, Yakir Berchenko
arXiv:2607. 00460v1 Announce Type: cross Abstract: Predicting complex spatiotemporal dynamics in physical processes often demands computationally expensive numerical methods or data-driven neural networks that suffer from high training costs, error accumulation, and limited generalizability to unseen parameters.
By Xin-Yang Liu, Xiantao Fan, Jian-Xun Wang
Invariant learning seeks representations that remain predictive across environments, yet the behavior of its objectives along the regularization path is often opaque. We address this objective-behavior gap by viewing representation learning as multimode magnetization and deriving, from concrete invariant-learning objectives, a Landau-type effective free energy whose low-order coefficients form objective signatures and induce distinct regularization phenotypes.
arXiv:2607. 03039v1 Announce Type: new Abstract: Neural networks are increasingly used to infer hidden physical structure from dynamical observations, yet it remains unclear whether their out-of-distribution performance reflects transferable physical rule learning.
By Yuan-Bin Zhu, Shuang Qiao, Shi-Ju Ran
arXiv:2607. 09801v1 Announce Type: new Abstract: Governing equations provide compact descriptions of physical systems, yet the variables in which they are simple are often hidden in high-dimensional measurements.
By Yi Zhu, Su Chen, Xiaojun Li, Xiuli Du
arXiv:2606. 16900v1 Announce Type: new Abstract: Physical systems often exhibit heterogeneous mechanisms, where rapidly evolving dynamics coexist with persistent structures.
By Hao Tang, Yuechen Duan, Jiongyu Zhu, Zimeng Feng, Hao Li, Chao Li