Expected free energy as an information constraint on the Bethe Lagrangian
arXiv:2608. 17167v1 Announce Type: cross Abstract: Active inference selects actions by minimising an expected free energy functional over predicted futures.
arXiv:2608. 15922v1 Announce Type: cross Abstract: We show that the natural-gradient stationary condition of variational inference has an edge-local form on a Forney-style factor graph.
arXiv:2608. 17167v1 Announce Type: cross Abstract: Active inference selects actions by minimising an expected free energy functional over predicted futures.
arXiv:2606. 06344v1 Announce Type: new Abstract: Probabilistic inference over spatially embedded variables requires beliefs that respect $SE(3)$ symmetry, yet existing equivariant networks produce only scalars and vectors -- not the rank-2 precision tensors needed for anisotropic uncertainty, and single-component messages collapse multi-modal energy landscapes to physically meaningless averages.
arXiv:2607. 13612v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) are the dominant design for latent world models, yet they are usually justified by empirical performance rather than a normative principle.
arXiv:2512. 24780v2 Announce Type: replace Abstract: Neural networks trained with standard objectives exhibit behaviors characteristic of probabilistic inference: soft clustering, prototype specialization, and Bayesian uncertainty tracking.
arXiv:2602. 04272v2 Announce Type: replace-cross Abstract: The Importance-Weighted Evidence Lower Bound (IW-ELBO) has emerged as an effective objective for variational inference (VI), tightening the standard ELBO and mitigating the mode-seeking behaviour.
arXiv:2608. 03967v1 Announce Type: cross Abstract: Generative Flow Networks (GFlowNets) have emerged as a flexible framework for amortised inference over discrete and mixed discrete-continuous objects, requiring only an unnormalised target density specified through a reward.
arXiv:2605. 30253v2 Announce Type: replace-cross Abstract: We study the contraction in Wasserstein distance of the coordinate ascent variational inference algorithm.
arXiv:2607. 10067v1 Announce Type: new Abstract: While autoregressive models optimize the exact data likelihood via the chain rule, diffusion models are typically trained with denoising objectives.
arXiv:2605. 08446v3 Announce Type: replace Abstract: Bayesian neural networks are typically trained against the evidence lower bound (ELBO), whose Jensen gap closes only when the variational posterior is exact.
arXiv:2604. 21407v2 Announce Type: replace Abstract: When approximating an intractable density via variational inference (VI) the variational family is typically chosen as a simple parametric family that very likely does not contain the target.
arXiv:2607. 17674v1 Announce Type: cross Abstract: A language model $p_\theta(y \mid x)$ trained on reasoning tasks learns to solve problems via multiple distinct strategies, yet these strategies are implicit and entangled within the model's response distribution.
arXiv:2606. 21585v4 Announce Type: replace Abstract: A finite agent, a machine's digital twin or any bounded reasoner, infers a fixed and noisy world through finite sensors, so its coherent output is a belief: a probability density over states (the Bayes posterior).