arXiv:2605. 08446v3 Announce Type: replace Abstract: Bayesian neural networks are typically trained against the evidence lower bound (ELBO), whose Jensen gap closes only when the variational posterior is exact.
By Pavel Prochazka
arXiv:2608.30720v1 Announce Type: new
Abstract: Representational similarity is foundational to analyses of deep networks, yet distances between point-valued representations are not intrinsically tied...
By Kieran Murphy
arXiv:2608. 03967v1 Announce Type: cross Abstract: Generative Flow Networks (GFlowNets) have emerged as a flexible framework for amortised inference over discrete and mixed discrete-continuous objects, requiring only an unnormalised target density specified through a reward.
By Yordan Raykov, Rodrigo Veiga
arXiv:2607. 24583v1 Announce Type: new Abstract: Large scale Bayesian nonparametrics (BNP) learner such as Stochastic Variational Inference (SVI) can handle datasets with large class number and large training size at fractional cost.
By Kart-Leong Lim
arXiv:2607. 08843v1 Announce Type: new Abstract: In artificial and biological neural networks, concepts are often encoded as consistent linear directions in representation space.
By William W. Yang, Andrew M. Saxe, Peter E. Latham
arXiv:2607. 09967v1 Announce Type: cross Abstract: Many neural networks operations have a multiplicative nature rather than additive: halving or doubling a norm are analogous relatively but require unequal optimization distances when taking linear steps.
By Ethan Smith
arXiv:2412. 08951v3 Announce Type: replace Abstract: Scalable algorithms of posterior approximation allow Bayesian nonparametrics such as Dirichlet process mixture to scale up to larger dataset at fractional cost.
By Kart-Leong Lim, Xudong Jiang
arXiv:2602. 22600v2 Announce Type: replace-cross Abstract: Training selects for behavior, not circuitry: many weight configurations can implement the same function.
By Joshua S. Schiffman
The paper introduces a data‑free learning objective called relative trajectory balance for training diffusion models to sample from a posterior defined by a diffusion prior and an arbitrary black‑box constraint or likelihood. It proves asymptotic correctness of this objective and demonstrates its use across vision, language, and multimodal tasks, including classifier guidance, language infilling, and text‑to‑image generation. Additionally, the method is applied to continuous control with a score‑based behavior prior, achieving state‑of‑the‑art results in offline reinforcement learning.
By Siddarth Venkatraman, Moksh Jain, Luca Scimeca, Minsu Kim, Marcin Sendera, Mohsin Hasan, Luke Rowe, Sarthak Mittal, Pablo Lemos, Emmanuel Bengio, Alexandre Adam, Jarrid Rector-Brooks, Yoshua Bengio, Glen Berseth, Esmeralda S. Whitammer
arXiv:2601. 07944v2 Announce Type: replace-cross Abstract: Since the turn of the century, approximate Bayesian inference has steadily evolved as new computational techniques have been incorporated to handle increasingly complex, large-scale predictive problems.
By Roy Shivam Ram Shreshtth, Arnab Hazra, Gourab Mukherjee
arXiv:2606. 16028v1 Announce Type: new Abstract: Modern deep learning architectures are increasingly multi-task and multi-modal, using a pretrained foundation model combined with task-specific, fine-tuned models.
By Thomas Dittrich, Oliver Potocki, Philipp Grohs
arXiv:2605. 08475v3 Announce Type: replace-cross Abstract: In this paper, we study in-context kernel ridge regression (KRR) with Gaussian kernels and show, both theoretically and empirically, that a standard softmax-attention transformer can approximate the KRR predictor during its forward pass.
By Mingsong Yan, Dongyang Li, Charles Kulick, Sui Tang