The paper introduces a generalized score matching objective for parameter estimation on convex subsets of ρ^d, derived from Minimum Probability Flow learning. It shows that this objective is a proper local scoring rule of second order, ensuring recovery of the true density when minimized, and proves convexity and consistency for exponential family models under standard conditions. Experiments demonstrate the method’s effectiveness on constrained domains where the partition function is intractable, including a generative modeling use‑case.
By Nishanth Shetty, Saisuchith Mahajan, Chandra Sekhar Seelamantula
FedFIbOS introduces a Fisher‑importance based criterion for selecting submodel parameters in heterogeneous federated learning, addressing the lack of theoretical justification in prior heuristic methods. By deriving a Fisher‑weighted quadratic masking surrogate and showing that the raw Fisher top‑k rule satisfies this surrogate under a Fisher‑dominant ranking condition, the method preserves convergence guarantees while efficiently estimating Fisher scores from squared gradients. Experiments on CIFAR‑10, CIFAR‑100, and AGNews demonstrate that FedFIbOS outperforms state‑of‑the‑art approaches by roughly 10% in accuracy, especially under strong non‑IID heterogeneity.
By Yasmeen Afzal, Jeremiah D. Deng, Haibo Zhang
arXiv:2504.05161v2 Announce Type: replace-cross
Abstract: Score estimation is the backbone of score-based generative models (SGMs), especially denoising diffusion probabilistic models (DDPMs). A key...
By Sinho Chewi, Alkis Kalavasis, Anay Mehrotra, Omar Montasser
arXiv:2606. 19587v1 Announce Type: cross Abstract: We propose a scalable method for training prediction (machine learning) models in the predict-then-optimize paradigm, where model outputs serve as coefficients for a subsequent linear optimization task.
By Beichen Wan, Mo Liu
The paper investigates Bayesian optimization using information geometry, deriving a local sensitivity tensor from the Fisher information metric that bounds the gradient of reparameterizable acquisition functions. This framework explains vanishing-gradient issues in high-dimensional settings and unifies heuristics like RAASP and dimension-scaled lengthscales. Leveraging this insight, the authors introduce FITR, a trust‑region BO method that replaces lengthscale scaling with local pullback‑Fisher weights, achieving competitive performance on GP benchmarks and extending naturally to non‑isotropic surrogates.
By Saksham Kiroriwal, Julius Pfrommer, J\"urgen Beyerer
arXiv:2606. 19876v1 Announce Type: new Abstract: The score matching problem is a central training objective in modern generative modeling, diffusion models, fitting unnormalized statistical models, and inverse problems.
By Alexander Tyurin
arXiv:2606. 23838v1 Announce Type: new Abstract: When two or more parameters or labels produce similar data, they are degenerate, or hard to distinguish.
By T. Lucas Makinen, Deaglan J. Bartlett, Niall Jeffrey, Benjamin D. Wandelt
arXiv:2607. 03871v1 Announce Type: new Abstract: Minimum maximum mean discrepancy (MMD) estimation has emerged as a robust and likelihood-free alternative to maximum likelihood estimation for parameter estimation.
By Sophia Seulkee Kang, Louis Sharrock, Xiaoyuan Cheng, Fran\c{c}ois-Xavier Briol, Zonghao Chen
The paper introduces a new variational inference framework that uses tangent transformations to handle strongly super‑Gaussian likelihoods across a wide range of probability models. By constructing tangent minorants of the log‑likelihood through convex duality, the method achieves conjugacy with Gaussian priors, enabling tractable inference where traditional approaches struggle. The authors provide algorithmic convergence guarantees and near‑parametric risk bounds, and demonstrate superior scalability and accuracy on both simulated and real‑world datasets compared to existing variational algorithms.
By Somjit Roy, Pritam Dey, Debdeep Pati, Bani K. Mallick
arXiv:2510. 24561v3 Announce Type: replace-cross Abstract: LoRA has become a widely adopted method for PEFT, and its initialization methods have attracted increasing attention.
By Qingyue Zhang, Chang Chu, Tianren Peng, Qi Li, Xiangyang Luo, Zhihao Jiang, Shao-Lun Huang
arXiv:2508. 03636v3 Announce Type: replace-cross Abstract: We propose a Likelihood Matching approach for training diffusion models by first establishing an equivalence between the likelihood of the target data distribution and a likelihood along the sample path of the reverse diffusion.
By Lei Qian, Wu Su, Yanqi Huang, Song Xi Chen
The paper introduces Gradient-based Sample Selection Bayesian Optimization (GSSBO), a method that builds the Gaussian process surrogate on a strategically chosen subset of samples rather than the full dataset. By using gradient information to eliminate redundant points while keeping diversity and representativeness, GSSBO achieves sublinear regret bounds and reduces the cubic computational cost of standard BO. Experiments on synthetic and real-world tasks show that this approach maintains comparable optimization performance while significantly cutting GP fitting time and resource usage.
By Qiyu Wei, Haowei Wang, Zirui Cao, Songhao Wang, Richard Allmendinger, Mauricio A \'Alvarez