arXiv AI

Scaling Up Thermodynamic AI Models

arXiv:2607. 00170v1 Announce Type: cross Abstract: Thermodynamic computing devices based on the Ising model show great promise for low-power AI inference and edge computing, but scalable methods for training large models for such hardware remain limited.

Hugging Face Trending Papers
Jun 30

Scaling Up Thermodynamic AI Models

Thermodynamic computing devices based on the Ising model show great promise for low-power AI inference and edge computing, but scalable methods for training large models for such hardware remain limited. Prior theory shows that the time-averaged behavior of high-temperature Gibbs-sampled Ising systems can implement feed-forward neural inference.

arXiv AI
Jun 9

Optimizing Energy-based Neural Network Training with Coherent Ising Machine

arXiv:2606. 09117v1 Announce Type: cross Abstract: While Ising machines serve as advanced physical solvers for the Ising model,enabling applications in combinatorial optimization and neural network training,their scalability for large-scale neural networks remains constrained by hardware connectivity limitations and suboptimal training methodologies.

By Chen-Rui Fan, Bo Lu, Zhi-Hong Zhang, Run-Qing Zhang, Jing-Wei Wen, Chuan Wang
Hugging Face Trending Papers
Jul 29

Equilibrium Training of Energy-Based Models with Parallel Trajectory Tempering

Energy-Based Models (EBMs) provide an interpretable framework for generative modeling of scientific data, but poor Markov Chain Monte Carlo mixing often limits their reliability. We introduce a training algorithm based on Parallel Trajectory Tempering (PTT), which exploits the continuity of the optimization path to maintain equilibrium sampling throughout learning.

arXiv Machine Learning
Aug 27

Thermodynamic cost of inference and learning in physical neural networks

The paper investigates the thermodynamic cost of inference and learning in physical neural networks. It shows that quasi‑static inference requires no work, while finite‑speed inference incurs work bounded by the Wasserstein‑2 distance between thermal states, roughly $k_B T$ per dimension of the widest layer. Learning, however, has an irreducible cost of a few $k_B T$ per parameter, independent of speed, indicating that memory dominates the thermodynamic price.

By Alexei V. Tkachenko
arXiv Machine Learning
Aug 19

Nonlocal Transition Kernel for Efficient Learning of Restricted Boltzmann Machines

The paper introduces a new transition kernel for Restricted Boltzmann Machines that operates over the sequence of models used in Deep Tempering. This kernel employs a round‑trip structure, allowing nonlocal moves in a single transition while keeping the RBM sequence unchanged. Experiments demonstrate that it achieves higher sampling quality with fewer transitions than both blocked Gibbs sampling and Deep Tempering, and it stabilizes learning by reducing training failures.

By Kaiji Sekimoto, Muneki Yasuda
arXiv Machine Learning
Jul 7

Out-of-distribution Neural Inference in Dynamical Ising Models

arXiv:2607. 03039v1 Announce Type: new Abstract: Neural networks are increasingly used to infer hidden physical structure from dynamical observations, yet it remains unclear whether their out-of-distribution performance reflects transferable physical rule learning.

By Yuan-Bin Zhu, Shuang Qiao, Shi-Ju Ran
arXiv AI
Sep 15

IsingFormer: Augmenting Parallel Tempering With Learned Proposals

The paper introduces IsingFormer, a Transformer model trained on long‑run MCMC configurations, which provides global proposal moves for Parallel Tempering (PT). By integrating these learned proposals into PT—forming Transformer‑Augmented Parallel Tempering (TAPT)—the authors demonstrate lower residual energies on 3D spin‑glass instances and improved efficiency on integer factorization tasks. A scaling study shows TAPT reduces the time‑to‑solution exponent by about 33% compared to standard PT across tested problem sizes.

By Saleh Bunaiyan, Corentin Delacour, Shuvro Chowdhury, Kyle Lee, Abdelrahman S. Abdelrahman, Kerem Y. Camsari
arXiv Machine Learning
Aug 31

Node-wise Feature Encoding for Neural Performance Prediction

FeatureFormer is a neural performance predictor that adds explicit node-wise encodings of FLOPs, parameter counts, and memory proxies to a gated graph attention architecture. It is designed to improve latency and energy prediction for neural networks on edge devices, addressing the limitation of existing GNN and transformer predictors that largely ignore node-level computational cost. The authors also introduce NNEQ, a large-scale energy consumption dataset, and show through extensive experiments that FeatureFormer achieves state‑of‑the‑art performance across both metrics, including challenging out‑of‑domain settings, while the encoding can broadly enhance existing predictors with negligible overhead.

By Matthew Grenier, William Hammer, Andrew Heuer, Nikhil Krishna, Yi Wang, Ramtin Zand
arXiv Machine Learning
Sep 16

Training Energy-Based Models with Non-MCMC Samplers and Efficient Temperature Estimation

The paper introduces Langevin simulated bifurcation (LSB), a fast, parallel Boltzmann sampler that matches the accuracy of sequential MCMC methods. It also proposes conditional expectation matching (CEM), an efficient technique for estimating the effective temperature of samples from energy‑based models with conditional independence. Building on these, the authors develop sampler adaptive learning (SAL), which adjusts the model temperature to align with the distribution produced by LSB, enabling efficient training of semi‑restricted Boltzmann machines (SRBMs) and outperforming conventional methods on synthetic spin‑glass datasets.

By Kentaro Kubo, Hayato Goto