Thermodynamic Limits of Physical Intelligence
arXiv:2602. 05463v2 Announce Type: replace-cross Abstract: Modern AI systems achieve remarkable capabilities at the cost of substantial energy consumption.
The paper investigates the thermodynamic cost of inference and learning in physical neural networks. It shows that quasi‑static inference requires no work, while finite‑speed inference incurs work bounded by the Wasserstein‑2 distance between thermal states, roughly $k_B T$ per dimension of the widest layer. Learning, however, has an irreducible cost of a few $k_B T$ per parameter, independent of speed, indicating that memory dominates the thermodynamic price.
arXiv:2602. 05463v2 Announce Type: replace-cross Abstract: Modern AI systems achieve remarkable capabilities at the cost of substantial energy consumption.
arXiv:2607. 27000v1 Announce Type: cross Abstract: Optimization in non-convex neural network models is strongly influenced by the geometry of the solution space: sparse, isolated, point-like clusters are typically algorithmically inaccessible, whereas wide and flat regions can be found efficiently despite being relatively rare.
arXiv:2607. 00170v1 Announce Type: cross Abstract: Thermodynamic computing devices based on the Ising model show great promise for low-power AI inference and edge computing, but scalable methods for training large models for such hardware remain limited.
Thermodynamic computing devices based on the Ising model show great promise for low-power AI inference and edge computing, but scalable methods for training large models for such hardware remain limited. Prior theory shows that the time-averaged behavior of high-temperature Gibbs-sampled Ising systems can implement feed-forward neural inference.
arXiv:2608. 16080v1 Announce Type: new Abstract: Thermal-aware optimization of multi-die 3D integrated circuits evaluates many designs, each a costly heat-equation solve.
arXiv:2608. 00097v1 Announce Type: cross Abstract: Physical learning rules such as equilibrium propagation (EP), coupled learning (CL), and adjoint coupled learning (AL) train resistive networks through local measurements.
arXiv:2606. 08343v1 Announce Type: new Abstract: We introduce GENERIC-FNO, the first neural operator to embed the full GENERIC (metriplectic) structure of nonequilibrium thermodynamics -- reversible, energy-conserving dynamics and irreversible, entropy-producing dynamics coupled through the degeneracy conditions -- directly in function space.
arXiv:2602. 03670v2 Announce Type: replace-cross Abstract: Equilibrium Propagation (EP) is a physics-inspired learning algorithm that uses stationary states of a dynamical system both for inference and learning.
arXiv:2608. 01357v1 Announce Type: new Abstract: Traditional approximation theory measures convergence rates in terms of the number of parameters or degrees of freedom.
arXiv:2606. 29519v1 Announce Type: new Abstract: Long-range learning is hard for recurrent networks trained with stochastic gradient descent, because the influence of a past input fades with the lag $\ell$, and if it fades too fast the dependence cannot be learned from finite data.
Long-range learning is hard for recurrent networks trained with stochastic gradient descent, because the influence of a past input fades with the lag $\ell$, and if it fades too fast the dependence cannot be learned from finite data. This fade is captured by an envelope $f(\ell)$.
The paper investigates how temperature affects analog deep neural network (DNN) inference, focusing on both stochastic and systematic non‑idealities in analog hardware. Experiments show that temperature‑induced performance loss is mainly driven by systematic errors rather than random noise. The study evaluates various mitigation techniques, finding that noise‑aware training and temperature‑aware calibration—especially hardware‑in‑the‑loop training—best preserve inference accuracy across different thermal conditions.