arXiv Machine Learning

The limits of interpretability in multiple linear regression

arXiv:2606. 16013v1 Announce Type: cross Abstract: Interpreting machine-learning models has attracted increasing attention, particularly in the physical sciences, where one often seeks to understand the underlying mechanisms rather than merely make predictions.

Hugging Face Trending Papers
Jun 10

Sparse probes and murky physics: a case study of interpretability challenges in a foundation model for continuum dynamics

Generative AI emulators are increasingly used in scientific domains where we already have strong theory, benchmarks, and physical intuition. This raises a central evaluation and interpretability question: when a foundation-style model can reproduce known continuum dynamics, what internal mechanism supports that behavior, is the internal behaviour consistent with known physics, and how does it relate to where the emulator succeeds or fails?

arXiv Machine Learning
1d ago

Stable and Counterfactually Robust Physical World Models from Imposed Structure and Learned Physics

The paper introduces a world model that learns to predict the evolution of physical systems while respecting key physical principles. By hard‑coding a general structure—generating dynamics from the gradient of a learned energy via a fixed reversible operator and imposing constraints on energy, dissipation, and interventions—the model achieves second‑law compatible dissipation, accurate responses to parameter changes, long‑term stability, and robustness to disturbances. Experiments on an electromagnetic cavity, a particle‑in‑cell grid, and shallow‑water fluid demonstrate that the model can recover accurate constitutive functions, distinguish conserving from dissipating regimes, and transfer learned physics to unseen conditions, outperforming unconstrained models.

By Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling
arXiv Machine Learning
Sep 23

A Spectral Theory of Grokking: Weight Decay induces Feature Learning

The paper presents a spectral theory explaining the phenomenon of grokking, where an initial fit to training data is followed by a delayed improvement in generalization. It shows that for homogeneous networks trained with squared loss and L₂ weight decay, residuals after memorization influence the neural tangent kernel (NTK) dynamics, leading to a transition from lazy to rich learning. The theory predicts that grokking timescales depend on the product of learning rate and weight decay, and that stronger decay can halt fitting, with empirical validation on modular addition tasks using MLPs and Transformers.

By Lenz Pracher, Pascal de Jong, Oskar Lieshaus, Alan Jeffares, Steffen Rulands