arXiv:2605. 19178v2 Announce Type: replace-cross Abstract: The great success of neural networks primarily arises from the presence of the large number of weight parameters combined with nonlinearities in the input-output relationship of single neurons.
By Giovanni di Sarra, Yasser Roudi
arXiv:2606. 31110v1 Announce Type: new Abstract: Artificial neural networks (NNs) and machine learning (ML) algorithms are poorly understood from a theoretical perspective, which makes it difficult to fully realize their potential and overcome their weaknesses.
By Robin Theriault
arXiv:2505. 11635v2 Announce Type: cross Abstract: Many real-world tasks, from associative memory to symbolic reasoning, benefit from discrete, structured representations that standard continuous latent models can struggle to express.
By Nikhil Kapasi, Mohamed Elfouly, William Whitehead, Luke Theogarajan
arXiv:2608.30978v1 Announce Type: new
Abstract: Modularity in deep neural networks has been proposed as a means of improving both interpretability and training by promoting disentangled representatio...
By Baptiste Rossigneux, Karim Haroun
arXiv:2512. 24780v2 Announce Type: replace Abstract: Neural networks trained with standard objectives exhibit behaviors characteristic of probabilistic inference: soft clustering, prototype specialization, and Bayesian uncertainty tracking.
By Alan Oursland
arXiv:2510.13210v2 Announce Type: replace
Abstract: We compare Ising ({-1, +1}) and QUBO ({0, 1}) encodings for Boltzmann machine learning under controlled protocols that fix the sampler, optimizer,...
By Yasushi Hasegawa, Masayuki Ohzeki
arXiv:2505. 11702v3 Announce Type: replace Abstract: This work develops a framework for post-training augmentation invariance, in which our goal is to add invariance properties to a pretrained network without altering its behavior on the original, non-augmented input distribution.
By Keenan Eikenberry, Lizuo Liu, Yoonsang Lee
The paper introduces a new transition kernel for Restricted Boltzmann Machines that operates over the sequence of models used in Deep Tempering. This kernel employs a round‑trip structure, allowing nonlocal moves in a single transition while keeping the RBM sequence unchanged. Experiments demonstrate that it achieves higher sampling quality with fewer transitions than both blocked Gibbs sampling and Deep Tempering, and it stabilizes learning by reducing training failures.
By Kaiji Sekimoto, Muneki Yasuda
arXiv:2509. 04899v4 Announce Type: replace-cross Abstract: Restricted Boltzmann machines (RBMs) are energy-based models originating from statistical physics, in which hidden units mediate the probability distribution of high-dimensional visible configurations.
By Mutsumi Kobayashi, Hiroshi Watanabe
arXiv:2606. 09725v1 Announce Type: new Abstract: Disentanglement, the separation of factors of variation in data using neural networks, remains a long-standing challenge in machine learning.
By Jhonny J. Velasquez Olivera, Christo K. Thomas, Walid Saad
arXiv:2610.07562v1 Announce Type: cross
Abstract: Learning an ensemble of GFlowNets to sample from a discrete target distribution has become a common approach for achieving better state space explora...
By Tiago da Silva, Amauri H. Souza, Salem Lahlou
arXiv:2609. 02672v1 Announce Type: cross Abstract: Hyper-Connections (HC) replace the single residual stream of a Transformer with $n$ parallel ones, mixing them at every layer with a learned $n \times n$ residual matrix.
By Haoqiang Guo, Xuyi Chen, Bo Ke, Yishu Lei, Ziyang Xu, Shikun Feng, Ximen, Wenhan Luo