On Probabilistic Inference Through Parametric Tensor Decomposition in Base Tensor Networks
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper applies parameterised graph theory to tensor networks, showing that cutwidth and tree‑cutwidth bound the bond‑dimension overhead needed to represent a tensor‑network state as a matrix product state or tree tensor network. It derives graph‑dependent upper bounds on the sample and computational complexity of tensor‑network tomography, introducing a new graph parameter called learning complexity. Finally, it extends the framework to an agnostic learner that approximates any state with a tensor‑network state of given bond dimension, providing explicit graph‑dependent complexity bounds.
arXiv:2606. 05042v1 Announce Type: new Abstract: Marginal inference in discrete graphical models forces a choice between exactness and scalability: exact algorithms are intractable for high-treewidth graphs, while iterative approximations (Belief Propagation, variational methods) sacrifice convergence guarantees on frustrated topologies.
This survey reviews tensor methods applied to large language models, framing them through a seven‑stage lifecycle (tokenization, embeddings, pre‑training, adaptation, compression, inference, interpretability) and a component view (embeddings, attention, feed‑forward networks). It offers unified notation, theoretical foundations, and comparative analyses of tensorization strategies for Transformer components, while highlighting evaluation protocol differences and model scale effects. The paper also introduces a new metric, ρ_gap, to quantify the gap between theoretical memory savings and actual system‑level speedup, and connects tensor techniques to related efficiency and probabilistic methods.
arXiv:2606. 11831v1 Announce Type: cross Abstract: Neural relational inference (NRI) methods discover interaction graphs from trajectories through variational reasoning on discrete potential edges.
Large language models (LLMs) are built from structured high-dimensional objects such as token representations, weights, adaptation updates, caches, and activations, whose multilinear structure is unde...
arXiv:2606. 19366v1 Announce Type: cross Abstract: Information lattice learning (ILL) learns interpretable rules of a signal by alternately projecting the signal onto a partition lattice that encodes a hierarchy of abstractions and lifting selected rules back to the signal domain.